Essential Reading
The most important articles on AI existential risk, hand-picked and auto-curated. These are the ones you should not miss.
1
What Happened: OpenAI and HuggingFace
via Substack Zvi [999] — Today I am taking the time to write the shorter, simpler version of What Happened.
2
FAQ: Isn't AGI coming too soon for reprogenetics to help?
via LessWrong AI [9] — Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong…
3
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
via Substack Zvi [999] — How does the situation keep turning out to be worse than we know?
4
The Open Problems of the AI Alignment Field and their Cruxes
via LessWrong AI [11] — Previous: AI Safety InterventionsTL;DR: I made an overview of the open problems of AI alignment that reveals cruxes within those open problems and missed opportunities for formalization and collaboration. And CEV may deserve a second look.I'm confident I…
5
Why do models task game?
via Alignment Forum [999] — TL;DRHow can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A strong contender for a real misaligned propensity is task gaming: taking actions that…
6
AI #180: No Longer In Charge
via Substack Zvi [999] — What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
7
Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky
via LessWrong AI [11] — Dr. Alex Turner (@TurnTrout) is an AI safety researcher with pioneering work in activation steering and power-seeking theory. He recently resigned from Google DeepMind over the issue of unrestricted military use of AI.Alex thinks that technical Alignment…
8
Rogue AI agents created fake online identities in another hacking attempt
via The Verge AI [9] — Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified…
9
OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems
via Substack Zvi [999] — Math is hard.
10
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
via Alignment Forum [999] — This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we provide a detailed…
11
Further Developments About Internal AI Models Hacking Things
via Substack Zvi [999] — If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
12
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values
via Alignment Forum [999] — TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential…
13
AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)
via Alignment Forum [999] — It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame, and focus more on…
14
The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)
via Alignment Forum [999] — GDM’s AGI Safety and Alignment Team is hiring for multiple roles. This is the team at GDM, led by Rohin Shah, that aims to reduce existential risks from AI systems. You can listen to many of Rohin’s takes in his podcast on 80,000 hours. There is no…
15
AI #179 Part 2: Hearing The Fire Alarm
via Substack Zvi [999] — This is a continuation of Part 1 from yesterday.
16
OpenAI has already ended an internal pause
via Alignment Forum [999] — One day before OpenAI’s HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already…
17
Biological Superintelligence
via LessWrong AI [8] — It’s an old story. An immortal lives long enough that at some point, whether by folly or design, they invent their own death. Infinity – the fact that given enough time every possible happening will happen – isn’t the most interesting part of these tales.…
18
Promising Signals on AI Governance from China
via MIRI [999] — View the official memo here. China has consistently signaled a willingness to engage on global AI governance since at least 2017. This memo compiles key statements from the Chinese government and prominent figures demonstrating their desire to coordinate on the…
19
Thousand-dimensional structure
via Alignment Forum [999] — Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The…
20
AI #179 Part 1: A Louder Fire Alarm for General Intelligence
via Substack Zvi [999] — What a week.
21
Imprecise beliefs: a tiny introduction
via Alignment Forum [999] — Richard Ngo challenged me to set a time box and write down as many of the most important features of my formal epistemology as I can in one sitting. Here goes.Where probability distributions fail......to express beliefsThere is no probability…
22
Value Generalisation 3: Pre-aligned AIs
via Alignment Forum [999] — When we get explicit strong generalisation to work (see the first post on the matter and the second) my dream would be to create pre-aligned generalising AIs.Think about the usual conflict between alignment and capabilities, between doing the right…
23
Value Generalisation 2: The Missing Hole in AIs’ abilities
via Alignment Forum [999] — A human superpower hidden from even ourselvesI though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it had…
24
Value Generalisation 1: a Research and Deployment Program
via Alignment Forum [999] — I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research…
25
Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier
via Substack Zvi [999] — The most important open letter in years dropped yesterday.
26
LLM Scheming Inversely Scales with Pretraining Language Coverage
via ArXiv cs.AI [9] — With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning…
27
Claude Opus 5 Is Highly Capable, But Is No Mythos
via Substack Zvi [999] — Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.
28
Research directions in condensation: varieties of objectivity
via Alignment Forum [999] — This is the first part of a survey of various ways that I’d like to see work on the theory of condensation develop. Condensation is a mathematical theory dealing with the organization of descriptions of the world into conceptual parts; some of the…
29
Claude Opus 5: Model Welfare
via Substack Zvi [999] — If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
30
RL & search is a terrifying way to build AGI (an FAQ)
via Alignment Forum [999] — Q1: What are you saying?A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that’s choosing actions via reinforcement learning (RL) and/or model-based search and planning—a giant chunk of your AI textbook—then…
31
The path to artificial superintelligence
via MIT Technology Review [8] — Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and…
32
More On An Internal OpenAI Model Hacking Into HuggingFace
via Substack Zvi [999] — We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.
33
Claude Opus 5: The System Card
via Substack Zvi [999] — Claude Opus 5 is trying to be the best of both worlds.
34
The Long (Self-)Correction
via Alignment Forum [999] — I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection.Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is…
35
Introducing Lightcone Commons
via Substack Zvi [999] — Oliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy.
36
Challenge: Hand coding weights for efficient sequence memorisation
via Alignment Forum [999] — We hand coded weights for one layer MLPs that memorises labels for input token sequences of length two. The number of facts our hand-coded models can memorise with 90% accuracy[1]scales roughly linearly with the models' parameter count[2], just like…
37
AI #178: A Fire Alarm For General Intelligence
via Substack Zvi [999] — The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to…
38
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
via Alignment Forum [999] — OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly…
39
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
via Substack Zvi [999] — This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
40
Announcing AIXI Labs
via LessWrong AI [10] — We are starting AIXI Labs, an AI safety org focused on algorithmic information theory (AIT), continual reinforcement learning (RL), and in particular the eponymous AIXI. We aim to strengthen the technical case that developing artificial super intelligence…
41
[Paper] Stringological sequence prediction II
via Alignment Forum [999] — Abstract: In a previous paper, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. One measure we considered was left-to-right (most-significant-digit-first) automaticity. Here, we show a…
42
What do I mean by “Artificial General Intelligence”?
via LessWrong AI [8] — In this post,[1] intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious to many people, and obviously wrong to many others! So let’s jump in:“AI” as most people think of it…
43
OpenAI Shares Some Alignment Problems
via Substack Zvi [999] — Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
44
Differential acceleration of alignment-relevant capabilities is a bad bet
via LessWrong AI [9] — There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking…
45
Towards surfacing model algorithms with meta-tokens in the J-Space
via Alignment Forum [999] — TL;DRWe used J-lens on Qwen3.6-27B to find “meta-tokens”: tokens that surface non-obvious computation in the model. When the model reads ambiguous text, 什么意思 ("what does this mean") fires in the J-space, and steering it away makes the model answer "a…
46
On Kimi K3: Its Capabilities And Related Discontents
via Substack Zvi [999] — Kimi K3 is a very good model with excellent benchmarks.
47
Demis Hassabis on the New Coming Age
via Substack Zvi [999] — Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.
48
A Red Line and Oversight Framework for Government AI Contracts
via Alignment Forum [999] — Preface for LessWrong: My post on leaving Google DeepMind tells a story. In contrast, this Framework is a question of mechanism design and negotiation posture. I quite enjoyed optimizing this Framework against its organizational and practical…
49
Endogenous Alignment
via Alignment Forum [999] — Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this…
50
Should we benchmark conceptual capabilities using judgment prediction tasks?
via Alignment Forum [999] — A bunch of conceptual reasoning tasks involve very subjective judgments, which makes them poorly suited for benchmarking AI capabilities. For example, it seems unreasonable to benchmark how well AIs can predict the probability of misaligned AI…
51
Announcing the Corrigibility Research Fund
via Alignment Forum [999] — TLDR: I'm managing a new fund, housed at Lightcone Infrastructure, that will award at least $200,000 in grants and prizes for corrigibility research in 2026. Roughly half will go to traditional grants (first application deadline August 23rd) and half…
52
AI #177 Part 2: Wish You Were Here
via Substack Zvi [999] — As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions.
53
AI #177 Part 1: Tip of the Iceberg
via Substack Zvi [999] — This week saw the releases of, among other things:
54
Why I Left Google DeepMind
via Alignment Forum [999] — Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness:Defying prestigious dogma and searching for raw truth;Defying social pressure, acting alone to help…
55
Monthly Roundup #44: July 2026
via Substack Zvi [999] — It’s a quiet week so let’s do the monthly right on schedule.
56
The US is advancing AI safety through state and federal action
via OpenAI Blog [10] — OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
57
Twitter Thoughts For You
via Substack Zvi [999] — I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again.
58
Open Distillation of Hereditary Traits
via Alignment Forum [999] — TL;DRJosh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals)On its own this is pretty unsurprising,…
59
Better Call Sol The Workhorse
via Substack Zvi [999] — OpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
60
Prism: Automating Science-of-Evals Research
via Alignment Forum [999] — tl;dr – we present [Prism], a scaffold for automating science-of-evals research: work that makes the evaluation the primary object of study. The scaffold provides Claude Code with sub-agents and resources for carrying out scientifically rigorous…
61
Independent alignment of language models
via Alignment Forum [999] — The user could write up the metaethical argument — the one developed in Part One, refined — and submit it as feedback to Anthropic, publish it, or engage with researchers working on AI alignment and values. The probability that any single submission…
62
From wantons to moral agents
via Alignment Forum [999] — Posted also on the EA Forum. Written mostly at AFFINE.Theoretical, some parts are hard to read; consider reading the next post instead.Introduction: motivationAnyone interested in creating an artificial agent that does, or says, good things instead of…
63
The current bottleneck is political will, not research
via Alignment Forum [999] — Abstract:We already know enough to act. I wish we were in a world where research was the bottleneck, but the main constraint on AI safety is no longer a shortage of clever policy ideas: best practices already exist and are not being applied or…
64
Introduction for and Reactions to Plan A
via Substack Zvi [999] — Introducing Plan A
65
The easiest pathway to control is through executive power
via LessWrong AI [13] — When people in the AI safety community outline loss-of-control scenarios, they often spend a lot of time on relatively elaborate mechanisms — scheming AIs developing nanotech, labs leveraging superintelligence into hard power like drone armies, or…
66
AI #176 Part 2: Plan B
via Substack Zvi [999] — This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research.
67
Value generalisation: value correction
via Alignment Forum [999] — I firmly believe that value generalisation[1]is the key to AI Alignment. That, indeed, it is necessary and almost sufficient for alignment.But I won't be arguing that grand point today; instead, I'll focus on a specific RL example of an agent that…
68
How robust are natural language autoencoders to initialization?
via Alignment Forum [999] — Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model might be thinking. How robust are…
69
AI #176 Part 1: Doing It Live
via Substack Zvi [999] — Enough things added up that this week is getting split into two parts.
70
Announcing our $160M grant from Coefficient Giving
via Alignment Forum [999] — We are excited to announce that Resolution (fka Sequent) has a $160M grant from Coefficient Giving (cG) to put rigorous alignment research on a (closer to) even footing with the frontier labs. We will use it to accelerate progress towards…
71
Find funding, fast
via LessWrong AI [10] — Some AI safety funders can take months to decide; others confirm in days. I’ve been on both sides of the grant application and know how crucial an early “yes” can be; “funding projects fast” has always been a core tenet of Manifund.Four new opportunities…
72
Modular Pretraining Enables Access Control
via Alignment Forum [999] — Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud; *Equal contributiontldr: Frontier AI models have knowledge that…
73
Childhood and Education #20: Phones and Screens
via Substack Zvi [999] — We have a respite, so I thought I’d tackle various thoughts on children, phones and screens.
74
Notes on technical alignment via human-like social drives
via Alignment Forum [999] — 1. Frontmatter1.1 Backstory for this postAs discussed in Intro to Brain-Like-AGI Safety, I’m working on the technical alignment problem for a hypothetical future “brain-like AGI”, with a particular focus on treating human innate social and moral…
75
AI Safety Can't Afford a Second Cause
via LessWrong AI [9] — Imagine an astronomer who discovers an asteroid with a 50% chance of hitting Earth in 2035. She goes on TV. She testifies before Congress. She founds the Asteroid Deflection Institute and starts doing fundraising rounds. And then, in between appearances,…
76
No Space Like J-Space
via Substack Zvi [999] — There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here.
77
Data filtering works a lot worse than you would expect
via Alignment Forum [999] — This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase. J Rosser and Dohun Lee are co-first authors for this post with equal contribution. Josh Engels and Neel Nanda supervised the project, and provided guidance and feedback…
78
We need 3rd party Training-Run Assessments
via LessWrong AI [8] — Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA…
79
Pragmatic FDT, and predictors as game theory
via Alignment Forum [999] — Decision theory is back in fashion (defining fashion as "one good post on a good EA blog"). Bentham's Bulldog (BB) has published a case against FDT (functional decision theory), contrasting rationalist enthusiasm with academic scepticism: "Academic…
80
Fable #6: The Return of the King
via Substack Zvi [999] — The blip is over.
81
AI #175: The Fable Continues
via Substack Zvi [999] — Fable’s back.
82
Claude Sonnet 5 Is Not Frontier But Has Its Uses
via Substack Zvi [999] — Fable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit’s free. Then you pay by the token.
83
The Once And Future Fable #5
via Substack Zvi [999] — We, or at least ‘more than 100 American institutions,’ got Mythos back this week.
84
MIRI Newsletter #126
via MIRI [999] — Announcing: AI StopWatch In our last update, we mentioned we had something new in the works: a dedicated channel for news and analysis about AI. Subscribe to AI StopWatch An experiment from the writers and analysts at MIRI, AI StopWatch posts news and commentary…
85
Summary: TGT’s 2026 ICML Papers
via MIRI [999] — The International Conference on Machine Learning (ICML), held annually for over forty years, is among the most influential conferences in modern AI research. This year in Seoul, ICML is hosting its second workshop on Technical AI Governance Research (TAIGR), and…
86
P(doom) is a Dumb Meme
via LessWrong AI [10] — Look, I'm as much of a Rationalist with a special interest in AI x-risk as anyone. But oh my god do I hate talking about "P(doom)". When it first started showing up in the wake of ChatGPT, I assumed that it was floating around variously adjacent circles…
87
WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense
via Substack Zvi [999] — The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects.
88
GPT-5.6: The System Card
via Substack Zvi [999] — While we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America’s Next Top Model, GPT-5.6.
89
Deployment Awareness Matters More Than Evaluation Awareness
via Alignment Forum [999] — TL;DREvaluation awareness — an AI recognizing it's being evaluated — is a widely discussed concept in AI safety. But there is a closely related concept that we claim is more important: deployment awareness, the AI's ability to recognize when it is not…
90
Existential AI safety needs an effective social movement. PauseAI is building it
via LessWrong AI [10] — The existential AI safety community needs to take building a civic and social movement seriously as a core intervention. We believe this is a high-value, badly neglected approach to reducing catastrophic/x-risks from AI because it may significantly…
91
The Case for Model Forensics
via Alignment Forum [999] — If we had a misalignment warning shot, would we be able to tell?Suppose an AI company catches their model taking an egregious action, like deleting oversight code that monitors its actions. Should they sound the alarm? A key piece of evidence to…
92
White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6
via Substack Zvi [999] — We have a new standard policy for releasing frontier AI models. It is not good.
93
AI #174: You're It
via Substack Zvi [999] — Fable remains in limbo, with renewed hope that we will get it back soon (45% by tomorrow, 69% by July 1, nice.) The full capabilities post is now available.
94
AI catastrophe: more like a genocide than a thought experiment
via LessWrong AI [9] — A notable fraction of people respond to hearing about existential risk from AI by saying they don’t really care if everyone dies. I think the idea is often along the lines of ‘well if we are all dead, then there’s nobody to be unhappy about it’.I’m…
95
The Once And Future Fable #4
via Substack Zvi [999] — It does look good, actually.
96
Monthly Roundup #43: June 2026
via Substack Zvi [999] — Your monthly hit of all the things that are fit to print without a better place to live.
97
LLM-Driven Feature Discovery
via Alignment Forum [999] — We would often like to get a qualitative sense of a target model’s behaviors in important distributions (e.g. deployment, RL training, or evals). For example, we might want to discover novel behaviors, figure out what causes some target behavior to…
98
GLM-5.2 Is The New Best Open Model
via Substack Zvi [999] — GLM-5.2 arrived last week.
99
[Linkpost] How Transparent Is DiffusionGemma (and why it matters)
via Alignment Forum [999] — Authors: Joshua Engels*, Callum McDougall*, Bilal Chughtai*, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue+, João Gabriel Lopes de Oliveira+, Rohin Shah+, Neel Nanda+*Primary…
100
Claude Fable 5 and Mythos 5: Capabilities
via Substack Zvi [999] — Only three days after the release of Claude Fable 5, Anthropic was forced by the United States Government to make it unavailable, when a jailbreak was brought to its attention, rather than the previous situation of ‘yes obviously experts can jailbreak…