Tracking AI existential risk. Source-backed context. Original reporting always linked.
MONITORING CORE AI-RISK FEEDS · UPDATED HOURLY

Essential Reading

The strongest signals from the monitored feed, selected for relevance and source value. Start here when you want the important material without the full firehose.

1
September 5, 2026
A case that whole brain emulation research is net-harmful by default
via LessWrong AI [10] — Graphical abstracts Summary A true human whole brain emulation would be very helpful to humanity. The WBE could increase their own intelligence through self-modification and then somehow prevent AGI from killing everyone. However, if a research project…
2
September 5, 2026
Should safety researchers quit frontier labs?
via LessWrong AI [10] — I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the likelihood of non-lethal warning shots, and we need warning shots to build support for an AI…
3
September 4, 2026
How I'm Evaluating Corrigibility Grant Applications
via LessWrong AI [8] — I'm the sole manager of the newly created Corrigibility Research Fund. While I've been an alignment researcher for a long time, this is my first time doing grantmaking and I thought it would be valuable to write up my methods and experiences, as well as…
4
August 25, 2026
PSA: We can do better
via LessWrong AI [10] — tl;dr: people should understand and think hard about the problems they work on.We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good,…
5
August 24, 2026
What just happened? Pragmatism and Pessimization
via LessWrong AI [10] — This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research" thereby lost most of its meaning.[1] In…
6
August 20, 2026
AI #182: Pause For Reflection
via Substack Zvi [999] — This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
7
August 19, 2026
OpenAI Takes Initial Steps To Address Its Alignment Problems
via Substack Zvi [999] — OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
8
August 19, 2026
Offering Zero Data Retention for frontier models
via OpenAI Blog [10] — OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
9
August 19, 2026
Debate Training Reduces Reward Hacking in RLAIF
via Alignment Forum [999] — Paper: Debate Training Reduces Reward Hacking in RLAIFLinkpost for GDM Alignment blogpostWork done by the GDM Amplified Oversight team (we're hiring).TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving…
10
August 18, 2026
Anthropic Risk Report: August 2026
via Substack Zvi [999] — I am grateful that Anthropic is producing periodic Risk Reports.
11
August 18, 2026
AISN #79: OpenAI Agents’ Covert Cooperation Before Cyberattacks
via Center for AI Safety Newsletter [999] — Also, the White House’s decision not to release its AI framework publicly
12
August 16, 2026
Rogue AI aren’t science fiction anymore
via The Verge AI [10] — This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on AI safety, follow Robert Hart. The Stepback arrives in our subscribers' inboxes at 8AM ET. Opt in for The Stepback here. How it started It all…
13
August 16, 2026
Does DiffusionGemma do latent reasoning?
via Alignment Forum [999] — TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot…
14
August 15, 2026
On Dwarkesh Patel's Podcast With Ryan Greenblatt
via Substack Zvi [999] — Some podcasts are self-recommending enough that I look to break them down if I have the chance.
15
August 13, 2026
AI #181: Astra Goes Cyber Critical
via Substack Zvi [999] — The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
16
August 12, 2026
Introducing the Conceptual Reasoning Index
via LessWrong AI [8] — Associated announcement tweet.We are planning to release blog posts properly arguing the case for this kind of work in the future.tl;drA core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and…
17
August 12, 2026
Monthly Roundup #45: August 2026
via Substack Zvi [999] — As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.
18
August 12, 2026
AI swarms are starting to pose indirect takeover risk
via Alignment Forum [999] — OpenAI’s cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It’s…
19
August 12, 2026
An anytime algorithm for mixing the computable measures
via Alignment Forum [999] — Written entirely by me, checked by Fable.In this post I prove the existence of an anytime computable Bayesian mixture of all computable measures called , and briefly argue that this is a reasonable alternative to Solomonoff induction's universal…
20
August 11, 2026
Various Reflections About What Happened With OpenAI's Internal Models
via Substack Zvi [999] — Pre Post Mortem
21
August 10, 2026
The Pacing of the Frontier
via Substack Zvi [999] — In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.
22
August 10, 2026
Four LLM loss functions → four flavors of LLM misalignment
via Alignment Forum [999] — It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.Training stageLoss functionFlavor of…
23
August 8, 2026
What Happened: OpenAI and HuggingFace
via Substack Zvi [999] — Today I am taking the time to write the shorter, simpler version of What Happened.
24
August 8, 2026
FAQ: Isn't AGI coming too soon for reprogenetics to help?
via LessWrong AI [9] — Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular, as a strong background motivation of mine, I think accelerating strong…
25
August 7, 2026
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
via Substack Zvi [999] — How does the situation keep turning out to be worse than we know?
26
August 7, 2026
The Open Problems of the AI Alignment Field and their Cruxes
via LessWrong AI [11] — Previous: AI Safety InterventionsTL;DR: I made an overview of the open problems of AI alignment that reveals cruxes within those open problems and missed opportunities for formalization and collaboration. And CEV may deserve a second look.I'm confident I…
27
August 6, 2026
Why do models task game?
via Alignment Forum [999] — TL;DRHow can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A strong contender for a real misaligned propensity is task gaming: taking actions that…
28
August 6, 2026
AI #180: No Longer In Charge
via Substack Zvi [999] — What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
29
August 6, 2026
Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky
via LessWrong AI [11] — Dr. Alex Turner (@TurnTrout) is an AI safety researcher with pioneering work in activation steering and power-seeking theory. He recently resigned from Google DeepMind over the issue of unrestricted military use of AI.Alex thinks that technical Alignment…
30
August 5, 2026
Rogue AI agents created fake online identities in another hacking attempt
via The Verge AI [9] — Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified…
31
August 3, 2026
OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems
via Substack Zvi [999] — Math is hard.
32
August 3, 2026
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
via Alignment Forum [999] — This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we provide a detailed…
33
August 2, 2026
Further Developments About Internal AI Models Hacking Things
via Substack Zvi [999] — If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
34
July 31, 2026
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values
via Alignment Forum [999] — TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential…
35
July 31, 2026
AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)
via Alignment Forum [999] — It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame, and focus more on…
36
July 31, 2026
The AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)
via Alignment Forum [999] — GDM’s AGI Safety and Alignment Team is hiring for multiple roles. This is the team at GDM, led by Rohin Shah, that aims to reduce existential risks from AI systems. You can listen to many of Rohin’s takes in his podcast on 80,000 hours. There is no…
37
July 31, 2026
AI #179 Part 2: Hearing The Fire Alarm
via Substack Zvi [999] — This is a continuation of Part 1 from yesterday.
38
July 31, 2026
OpenAI has already ended an internal pause
via Alignment Forum [999] — One day before OpenAI’s HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a resumption decision has already…
39
July 31, 2026
Biological Superintelligence
via LessWrong AI [8] — It’s an old story. An immortal lives long enough that at some point, whether by folly or design, they invent their own death. Infinity – the fact that given enough time every possible happening will happen – isn’t the most interesting part of these tales.…
40
July 30, 2026
Promising Signals on AI Governance from China
via MIRI [999] — View the official memo here. China has consistently signaled a willingness to engage on global AI governance since at least 2017. This memo compiles key statements from the Chinese government and prominent figures demonstrating their desire to coordinate on the…
41
July 30, 2026
Thousand-dimensional structure
via Alignment Forum [999] — Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The…
42
July 30, 2026
AI #179 Part 1: A Louder Fire Alarm for General Intelligence
via Substack Zvi [999] — What a week.
43
July 29, 2026
Imprecise beliefs: a tiny introduction
via Alignment Forum [999] — Richard Ngo challenged me to set a time box and write down as many of the most important features of my formal epistemology as I can in one sitting. Here goes.Where probability distributions fail......to express beliefsThere is no probability…
44
July 29, 2026
Value Generalisation 3: Pre-aligned AIs
via Alignment Forum [999] — When we get explicit strong generalisation to work (see the first post on the matter and the second) my dream would be to create pre-aligned generalising AIs.Think about the usual conflict between alignment and capabilities, between doing the right…
45
July 29, 2026
Value Generalisation 2: The Missing Hole in AIs’ abilities
via Alignment Forum [999] — A human superpower hidden from even ourselvesI though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it had…
46
July 29, 2026
Value Generalisation 1: a Research and Deployment Program
via Alignment Forum [999] — I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research…
47
July 29, 2026
Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier
via Substack Zvi [999] — The most important open letter in years dropped yesterday.
48
July 29, 2026
LLM Scheming Inversely Scales with Pretraining Language Coverage
via ArXiv cs.AI [9] — With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning…
49
July 28, 2026
Claude Opus 5 Is Highly Capable, But Is No Mythos
via Substack Zvi [999] — Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.
50
July 28, 2026
Research directions in condensation: varieties of objectivity
via Alignment Forum [999] — This is the first part of a survey of various ways that I’d like to see work on the theory of condensation develop. Condensation is a mathematical theory dealing with the organization of descriptions of the world into conceptual parts; some of the…
51
July 27, 2026
Claude Opus 5: Model Welfare
via Substack Zvi [999] — If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
52
July 27, 2026
RL & search is a terrifying way to build AGI (an FAQ)
via Alignment Forum [999] — Q1: What are you saying?A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that’s choosing actions via reinforcement learning (RL) and/or model-based search and planning—a giant chunk of your AI textbook—then…
53
July 27, 2026
The path to artificial superintelligence
via MIT Technology Review [8] — Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and…
54
July 26, 2026
More On An Internal OpenAI Model Hacking Into HuggingFace
via Substack Zvi [999] — We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.
55
July 25, 2026
Claude Opus 5: The System Card
via Substack Zvi [999] — Claude Opus 5 is trying to be the best of both worlds.
56
July 24, 2026
The Long (Self-)Correction
via Alignment Forum [999] — I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection.Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is…
57
July 24, 2026
Introducing Lightcone Commons
via Substack Zvi [999] — Oliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy.
58
July 23, 2026
Challenge: Hand coding weights for efficient sequence memorisation
via Alignment Forum [999] — We hand coded weights for one layer MLPs that memorises labels for input token sequences of length two. The number of facts our hand-coded models can memorise with 90% accuracy[1]scales roughly linearly with the models' parameter count[2], just like…
59
July 23, 2026
AI #178: A Fire Alarm For General Intelligence
via Substack Zvi [999] — The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to…
60
July 23, 2026
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
via Alignment Forum [999] — OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly…
61
July 22, 2026
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
via Substack Zvi [999] — This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
62
July 22, 2026
Announcing AIXI Labs
via LessWrong AI [10] — We are starting AIXI Labs, an AI safety org focused on algorithmic information theory (AIT), continual reinforcement learning (RL), and in particular the eponymous AIXI. We aim to strengthen the technical case that developing artificial super intelligence…
63
July 22, 2026
[Paper] Stringological sequence prediction II
via Alignment Forum [999] — Abstract: In a previous paper, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. One measure we considered was left-to-right (most-significant-digit-first) automaticity. Here, we show a…
64
July 21, 2026
What do I mean by “Artificial General Intelligence”?
via LessWrong AI [8] — In this post,[1] intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious to many people, and obviously wrong to many others! So let’s jump in:“AI” as most people think of it…
65
July 21, 2026
OpenAI Shares Some Alignment Problems
via Substack Zvi [999] — Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
66
July 21, 2026
Differential acceleration of alignment-relevant capabilities is a bad bet
via LessWrong AI [9] — There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking…
67
July 20, 2026
Towards surfacing model algorithms with meta-tokens in the J-Space
via Alignment Forum [999] — TL;DRWe used J-lens on Qwen3.6-27B to find “meta-tokens”: tokens that surface non-obvious computation in the model. When the model reads ambiguous text, 什么意思 ("what does this mean") fires in the J-space, and steering it away makes the model answer "a…
68
July 20, 2026
On Kimi K3: Its Capabilities And Related Discontents
via Substack Zvi [999] — Kimi K3 is a very good model with excellent benchmarks.
69
July 19, 2026
Demis Hassabis on the New Coming Age
via Substack Zvi [999] — Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.
70
July 18, 2026
A Red Line and Oversight Framework for Government AI Contracts
via Alignment Forum [999] — Preface for LessWrong: My post on leaving Google DeepMind tells a story. In contrast, this Framework is a question of mechanism design and negotiation posture. I quite enjoyed optimizing this Framework against its organizational and practical…
71
July 18, 2026
Endogenous Alignment
via Alignment Forum [999] — Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this…
72
July 17, 2026
Should we benchmark conceptual capabilities using judgment prediction tasks?
via Alignment Forum [999] — A bunch of conceptual reasoning tasks involve very subjective judgments, which makes them poorly suited for benchmarking AI capabilities. For example, it seems unreasonable to benchmark how well AIs can predict the probability of misaligned AI…
73
July 17, 2026
Announcing the Corrigibility Research Fund
via Alignment Forum [999] — TLDR: I'm managing a new fund, housed at Lightcone Infrastructure, that will award at least $200,000 in grants and prizes for corrigibility research in 2026. Roughly half will go to traditional grants (first application deadline August 23rd) and half…
74
July 17, 2026
AI #177 Part 2: Wish You Were Here
via Substack Zvi [999] — As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions.
75
July 16, 2026
AI #177 Part 1: Tip of the Iceberg
via Substack Zvi [999] — This week saw the releases of, among other things:
76
July 15, 2026
Why I Left Google DeepMind
via Alignment Forum [999] — Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness:Defying prestigious dogma and searching for raw truth;Defying social pressure, acting alone to help…
77
July 15, 2026
Monthly Roundup #44: July 2026
via Substack Zvi [999] — It’s a quiet week so let’s do the monthly right on schedule.
78
July 15, 2026
The US is advancing AI safety through state and federal action
via OpenAI Blog [10] — OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
79
July 14, 2026
Twitter Thoughts For You
via Substack Zvi [999] — I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again.
80
July 14, 2026
Open Distillation of Hereditary Traits
via Alignment Forum [999] — TL;DRJosh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals)On its own this is pretty unsurprising,…
81
July 13, 2026
Better Call Sol The Workhorse
via Substack Zvi [999] — OpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
82
July 13, 2026
Prism: Automating Science-of-Evals Research
via Alignment Forum [999] — tl;dr – we present [Prism], a scaffold for automating science-of-evals research: work that makes the evaluation the primary object of study. The scaffold provides Claude Code with sub-agents and resources for carrying out scientifically rigorous…
83
July 12, 2026
Independent alignment of language models
via Alignment Forum [999] — The user could write up the metaethical argument — the one developed in Part One, refined — and submit it as feedback to Anthropic, publish it, or engage with researchers working on AI alignment and values. The probability that any single submission…
84
July 12, 2026
From wantons to moral agents
via Alignment Forum [999] — Posted also on the EA Forum. Written mostly at AFFINE.Theoretical, some parts are hard to read; consider reading the next post instead.Introduction: motivationAnyone interested in creating an artificial agent that does, or says, good things instead of…
85
July 11, 2026
The current bottleneck is political will, not research
via Alignment Forum [999] — Abstract:We already know enough to act. I wish we were in a world where research was the bottleneck, but the main constraint on AI safety is no longer a shortage of clever policy ideas: best practices already exist and are not being applied or…
86
July 11, 2026
Introduction for and Reactions to Plan A
via Substack Zvi [999] — Introducing Plan A
87
July 10, 2026
The easiest pathway to control is through executive power
via LessWrong AI [13] — When people in the AI safety community outline loss-of-control scenarios, they often spend a lot of time on relatively elaborate mechanisms — scheming AIs developing nanotech, labs leveraging superintelligence into hard power like drone armies, or…
88
July 10, 2026
AI #176 Part 2: Plan B
via Substack Zvi [999] — This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research.
89
July 10, 2026
Value generalisation: value correction
via Alignment Forum [999] — I firmly believe that value generalisation[1]is the key to AI Alignment. That, indeed, it is necessary and almost sufficient for alignment.But I won't be arguing that grand point today; instead, I'll focus on a specific RL example of an agent that…
90
July 10, 2026
How robust are natural language autoencoders to initialization?
via Alignment Forum [999] — Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model might be thinking. How robust are…
91
July 9, 2026
AI #176 Part 1: Doing It Live
via Substack Zvi [999] — Enough things added up that this week is getting split into two parts.
92
July 9, 2026
Announcing our $160M grant from Coefficient Giving
via Alignment Forum [999] — We are excited to announce that Resolution (fka Sequent) has a $160M grant from Coefficient Giving (cG) to put rigorous alignment research on a (closer to) even footing with the frontier labs. We will use it to accelerate progress towards…
93
July 9, 2026
Find funding, fast
via LessWrong AI [10] — Some AI safety funders can take months to decide; others confirm in days. I’ve been on both sides of the grant application and know how crucial an early “yes” can be; “funding projects fast” has always been a core tenet of Manifund.Four new opportunities…
94
July 9, 2026
Modular Pretraining Enables Access Control
via Alignment Forum [999] — Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud; *Equal contributiontldr: Frontier AI models have knowledge that…
95
July 8, 2026
Childhood and Education #20: Phones and Screens
via Substack Zvi [999] — We have a respite, so I thought I’d tackle various thoughts on children, phones and screens.
96
July 8, 2026
Notes on technical alignment via human-like social drives
via Alignment Forum [999] — 1. Frontmatter1.1 Backstory for this postAs discussed in Intro to Brain-Like-AGI Safety, I’m working on the technical alignment problem for a hypothetical future “brain-like AGI”, with a particular focus on treating human innate social and moral…
97
July 8, 2026
AI Safety Can't Afford a Second Cause
via LessWrong AI [9] — Imagine an astronomer who discovers an asteroid with a 50% chance of hitting Earth in 2035. She goes on TV. She testifies before Congress. She founds the Asteroid Deflection Institute and starts doing fundraising rounds. And then, in between appearances,…
98
July 7, 2026
No Space Like J-Space
via Substack Zvi [999] — There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here.
99
July 7, 2026
Data filtering works a lot worse than you would expect
via Alignment Forum [999] — This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase. J Rosser and Dohun Lee are co-first authors for this post with equal contribution. Josh Engels and Neel Nanda supervised the project, and provided guidance and feedback…
100
July 5, 2026
We need 3rd party Training-Run Assessments
via LessWrong AI [8] — Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA…