Tracking AI existential risk. Auto-aggregated headlines. Human-curated analysis.
AGGREGATING 47 SOURCES · UPDATED LIVE
Analysis
Zac Boring 4 months ago Analysis
What do we know about AI company employee giving?
via LessWrong AI [7] — Many Anthropic employees, especially, are sympathetic to AI safety and (will) have lots of money. This is something that is being talked about a lot (semi-)privately, but I haven't seen any public discussion of it. I find that striking. It seems like the…
Zac Boring 4 months ago Analysis
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
via LessWrong AI [3] — TL;DR We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked.…
Zac Boring 4 months ago Analysis
Interview with Steven Byrnes on His Mainline Takeoff Scenario
via LessWrong AI [9] — After using the latest version of Claude Code and being surprised how capable it's become while still behaving friendly and corrigibly, I wanted to reflect on how this new observation should update my world model and my P(Doom).So I reached out to Dr.…
Zac Boring 4 months ago Analysis
The case for AI safety capacity-building work
via LessWrong AI [7] — TL;DR:I think many of the marginal hires at larger organizations doing AI safety technical or policy work right now (including e.g. Apollo, Redwood, METR, RAND TASP, GovAI, Epoch, UKAISI, and Anthropic’s safety teams) would be capable of founding (or being…
Zac Boring 4 months ago Analysis
Claude Code, Claude Cowork and Codex #5
via Substack Zvi [999] — It feels good to get back to some of the fun stuff.
Zac Boring 4 months ago Analysis
Promoting enmity and bad vibes around AI safety
via LessWrong AI [9] — I've observed some people engaged in activities that I believe are promoting enmity in the course of their efforts to raise awareness about AI risk. To be frank, I think those activities are increasing AI risk, including but not limited to extinction risk.…
Zac Boring 4 months ago Analysis
Payorian cooperation is easy with Kripke frames
via LessWrong AI [3] — The context is MIRI's twist on Axelrod's Prisoner's Dilemma tournament. Axelrod's competitors were programs, facing each other in an iterated Prisoner's Dilemma. MIRI's tournament is a one-shot Prisoner's Dilemma, but the programs get to read their…
Zac Boring 4 months ago Analysis
Your Causal Variables Are Irreducibly Subjective
via LessWrong AI [7] — Mechanistic interpretability needs its own shoe leather era. Reproducing the labeling process will matter more than reproducing the Github. And who can blame us? Causal inference comes with an impressive toolkit: directed acyclic graphs, potential…
Zac Boring 4 months ago Analysis
Mox is the largest AI Safety community space in San Francisco. We're fundraising!
via LessWrong AI [5] — Summary: Mox is fundraising to maintain and grow AIS projects, build a compelling membership, and foster other impactful and delightful work. We're looking to raise $450k for 2026, and you can donate on Manifund!OverviewWho we areMox is SF’s largest AI…
Zac Boring 4 months ago Analysis
Thoughts on the Pause AI protest
via LessWrong AI [4] — On Saturday (Feb 28, 2026) I attended my first ever protest. It was jointly organized by PauseAI, Pull the Plug and a handful of other groups I forget. I have mixed feelings about it. To be clear about where I stand: I believe that AI labs are worryingly…
Zac Boring 5 months ago Analysis
Anthropic Officially, Arbitrarily and Capriciously Designated a Supply Chain Risk
via Substack Zvi [999] — Make no mistake about what is happening.
Zac Boring 5 months ago Analysis
The Elect
via LessWrong AI — I was different in Michael’s prison than I was outside, looking the way I did when we fell in love so long ago, in that time before we could change our forms. Stuck in some body that was not of my choosing? Does that seem strange to you? It was not like that…
Zac Boring 5 months ago Analysis
Shaping the exploration of the motivation-space matters for AI safety
via LessWrong AI [5] — SummaryWe argue that shaping RL exploration, and especially the exploration of the motivation-space, is understudied in AI safety and could be influential in mitigating risks. Several recent discussions hint in this direction — the entangled generalization…
Zac Boring 5 months ago Analysis
AI Safety Has 12 Months Left
via LessWrong AI — The past decade of technology has been defined by many wondering what the upper bound of power and influence is for an individual company. The core concern about AI labs is that the upper bound is infinite.[1]This has led investors to direct all of their mindshare towards deploying into AI, the tech
Zac Boring 5 months ago Analysis
Personality Self-Replicators
via LessWrong AI [5] — One-sentence summaryI describe the risk of personality self-replicators, the threat of OpenClaw-like agents managing spreading in hard-to-control ways. SummaryLLM agents like OpenClaw are defined by a small set of text files and are run by an open source framework which leverages LLMs
Zac Boring 5 months ago Analysis
AI #158: The Department of War
via Substack Zvi — This was the worst week I have had in quite a while, maybe ever.
Zac Boring 5 months ago Analysis
Gemini 3.1 Pro Aces Benchmarks, I Suppose
via Substack Zvi — I’ve been trying to find a slot for this one for a while.
Zac Boring 5 months ago Analysis
Mass Surveillance w/ LLMs is the Default Outcome. Contracts Won't Change That.
via LessWrong AI [3] — What's the best case scenario regarding OpenAI's contract w/ the Department of War (DoW)?We have access to the full contractIt's airtightOAI's engineers are on top of things in case the DoW breaks the contractThere's actual teeth for violationsBut even then, the DoW can simply switch vendors. Use Ge
Zac Boring 5 months ago Analysis
I Had Claude Read Every AI Safety Paper Since 2020, Here's the DB
via LessWrong AI — Click here if you just want to see the Database I made of all[1] AI safety papers written since 2020 and not read the methodology. To some extent the core idea here is to encode as much info from these papers into something small enough that an AI with a specific problem in mind can take in all
Zac Boring 5 months ago Analysis
An Alignment Journal: Coming Soon
via LessWrong AI [9] — tl;dr We’re incubating an academic journal for AI alignment: rapid peer-review of foundational Alignment research that the current publication ecosystem underserves. Key bets: paid attributed review, reviewer-written synthesis abstracts, and targeted automation. Contact us if…
Live Doom Meter
-- %
0% — We're fine 100% — GG
P(Doom) Scoreboard
0%25%50%75%100%
Loading estimates...