Analysis
What if Anthropic unilaterally paused capabilities development right now?
via LessWrong AI [6] — In their new post on recursive self-improvement, Anthropic argues that a pause in frontier AI development is needed, but unfortunately, they can't pause on their own, because of less cautious actors:We believe it would be good for the world to have the…
Preparing for Warning Shots to Catalyze International Cooperation on AGI Risks
via LessWrong AI [4] — SummaryThis is a write-up on preparing for warning shots to catalyze international cooperation on AGI risks, and the corollary list of projects one could pursue. We argue we must first (1) understand types of warning shots, then (2) prepare to catch them.…
Learnings from starting an AI safety research team
via LessWrong AI [9] — This post’s goal is to distill our takeaways from building a new research team over the past four months. We describe some context about our team, how it came about, and then describe the lessons learned.Since AI safety is becoming more and more…
OpenAI Offers A New Policy Blueprint
via Substack Zvi [999] — Right after a new Executive Order seems like an excellent time to offer OpenAI’s new document: Democratic Governance of Frontier AI: A Blueprint For A Federal Framework.
Rohin Shah on AGI Safety
via LessWrong AI [6] — Rohin Shah recently had an interview on 80000 hours on his views on AGI Safety and his work at Google DeepMind. I'm posting the transcript below to encourage further discussion. I think the interview is interesting though I disagree on a bunch of topics,…
Sixteen schemes for AI safety
via LessWrong AI [5] — These days, I often run across whippersnappers excited to do something for AI safety — but aren’t quite sure what. One of the fun things about the Future Fund era were the big lists of project ideas; as we enter a new era of crazy money sloshing around, it…
AI #171: False Flag
via Substack Zvi [999] — This was the week of Claude Opus 4.8.
Society Explained: a tool for efficiently exploring >100 theories of society
via LessWrong AI [3] — There are many competing theories of how society does and should function, from Karl Marx and Adam Smith to Steven Pinker and Eliezer Yudkowsky. These theories are often hard to understand - you may need to read an entire book (or dozens of articles) to…
Trump Signs Executive Order For AI Testing Prior To Frontier Model Releases
via Substack Zvi [999] — Last week we were expecting an Executive Order on Thursday.
China won’t win the AI race but would it be much worse if it did?
via LessWrong AI [4] — It seems to me accepted wisdom in the West that the US owned labs must “beat” the Chinese labs in the race for AGI/ASI. Even those who don’t think there will be a winner, that essentially the race is to see which country’s AI will kill/disempower us first,…
Why Even Experts Don’t Know What to Do About AI Risk
via LessWrong AI [9] — AI Safety veteran Holden Karnofsky thinks there’s a 49% chance his actions are making things worse.[1]In 2025, Jesse Clifton even stepped down as the executive director of the Center on Long-Term risk because of similar reasons.Even top AI Safety…
Claude Opus 4.8: Capabilities and Reactions
via Substack Zvi [999] — You need a lot of data points to understand a new model, and what you have.
"Contagious Humming" to Silence a Room
via LessWrong AI [4] — Often when running meetups you’ll have several lively conversations going at the same time. This is a great problem to have, but it can make it difficult to get everyone’s attention for announcements.Try using “Contagious Humming” when you need to silence…
NYT: Senator Sanders Proposes Gov't Take 50% Ownership of AI Labs
via LessWrong AI [4] — Quoting from Senator Bernie Sanders Op-Ed in the New York Times today:(...) I will soon be introducing the American A.I. Sovereign Wealth Fund Act. This legislation would give the public a direct ownership stake in the largest A.I. companies in our…
Opus 4.8 Part 2: Model Welfare
via Substack Zvi [999] — Everything impacts everything.
When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability
via LessWrong AI [7] — We've found a method that tells you:How functionally similar two neural networks are across ALL inputs,Computed solely from the weights (i.e. no data),Using a principled generalization of cosine similarity.There's only one catch: you have to use a tensor…
Announcing: Iliad's Fall 2026 Programs
via LessWrong AI [5] — The April 2026 Iliad Intensive cohort, at LISAIliad, an umbrella organization for applied math for AI alignment, is running several additional programs through the end of the year!Applications to all of them are now open, here. Applicants will be selected…
Claude Opus 4.8: The System Card
via Substack Zvi [999] — Only six weeks after Opus 4.7, we have Opus 4.8.
Developmental Cognitive Interpretability: A Research Agenda for Modelling Generalisation and Predicting Agent Behaviour
via LessWrong AI [3] — SummarySafe deployment of an AI system requires that we can make confident claims about its behaviour on out-of-distribution deployment inputs on the basis of only pre-deployment evaluations. One approach to making such claims is to take a cognitive…
How can the middle powers avoid getting trounced during the intelligence explosion? A plan.
via LessWrong AI [4] — This is an edited version of a LW shortform.Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government. This will massively boost US military power and make the US economically…
Live Doom Meter
--
%
0% — We're fine
100% — GG
The Doom Meter is a composite score derived from prediction markets and feed sentiment, updated daily.
70%
Prediction Markets
Weighted average of Manifold Markets questions on AI catastrophe, AGI timelines, expert surveys, and key figures. Direct doom indicators weighted higher than indirect capability markers.
30%
Feed Sentiment
Percentage of recent headlines containing high-alarm keywords (existential risk, catastrophe, extinction). Higher alarm density = higher score.
This is not a scientific estimate of existential risk. It is an opinionated, transparent signal — a vibes-based thermometer for AI doom discourse.
P(Doom) Scoreboard
0%25%50%75%100%
Loading estimates...