Tracking AI existential risk. Auto-aggregated headlines. Human-curated analysis.
AGGREGATING 47 SOURCES · UPDATED LIVE
Analysis
Zac Boring 2 months ago Analysis
Classifier Context Rot: Monitor Performance Degrades with Context Length
via LessWrong AI [3] — Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100K tokens.We show that when used as…
Zac Boring 2 months ago Analysis
Dating Roundup #12: Sex and Violence
via Substack Zvi [999] — No more burying the sex stuff under an avalanche of other stuff so no one notices.
Zac Boring 2 months ago Analysis
An Introduction to Exemplar Partitioning for Mechanistic Interpretability
via LessWrong AI [7] — Most of what we currently call "feature discovery" in language models is wrapped up in dictionary-learning methods like sparse autoencoders (SAEs) – which work, and which have been scaled to millions of features on frontier-scale models, but which bundle…
Zac Boring 2 months ago Analysis
A Year Late, Claude Finally Beats Pokémon
via LessWrong AI [3] — Credit: ClaudePlaysPokemon Elevator Shanty by KurukkooDisclaimer: like some previous posts in this series, this was not primarily written by me, but by a friend. I did substantial editing, however.ClaudePlaysPokemon feat. Opus 4.7 has finally beaten…
Zac Boring 2 months ago Analysis
The hard core of alignment (is robustifying RL)
via LessWrong AI [5] — Most technical AI safety work that I read seems to miss the mark, failing to make any progress on the hard part of the problem. I think this is a common sentiment, but there's less agreement about what exactly the hard part is? Characterizing this more…
Zac Boring 2 months ago Analysis
Monthly Roundup #42: May 2026
via Substack Zvi [999] — At least we probably won’t have another pandemic.
Zac Boring 2 months ago Analysis
Convergent Abstraction Hypothesis
via LessWrong AI [4] — Tl;drConvergent abstraction hypothesis posits abstractions are often convergent in the sense of convergent evolution: different cognitive systems converge on the same abstraction, when facing similar selection pressures and learning in similar…
Zac Boring 2 months ago Analysis
AI #168: Not Leading the Future
via Substack Zvi [999] — This is what a lull looks like at this point.
Zac Boring 2 months ago Analysis
Most "inner work" looks like entertainment.
via LessWrong AI [4] — Imagine you’re looking for a personal trainer. You open one trainer’s webpage and read their testimonials: “I had an experience tied for the most intense experiences of my life”; “They do it all with fun, care, and a sense of humour.” You notice that none…
Zac Boring 2 months ago Analysis
Cyber Lack of Security and AI Governance
via Substack Zvi [999] — The real recent story of AI has been the background work being done on Cybersecurity, as we process the Mythos Moment along with GPT-5.5, and figure out both how to patch the internet and what our new regulatory regime is going to look like.
Zac Boring 2 months ago Analysis
Voters are surprisingly open to talking about AI risk
via LessWrong AI [14] — TL;DR: Voters are now surprisingly open to talking about existential risk from AI. This seems to have changed in the last 6 months. When campaigning for AI safety-friendly politicians (e.g., Alex Bores), we should talk more about AI in general, and about…
Zac Boring 2 months ago Analysis
Childhood and Education #18: Do The Math
via Substack Zvi [999] — We did reading yesterday.
Zac Boring 2 months ago Analysis
The Iliad Intensive Course Materials
via LessWrong AI [5] — We are releasing the course materials of the Iliad Intensive, a new month-long and full-time AI Alignment course that runs in-person every second month. The course targets students with strong backgrounds in mathematics, physics, or theoretical computer…
Zac Boring 2 months ago Analysis
Childhood And Education #17: Is Our Children Reading
via Substack Zvi [999] — Reading is the most fundamental thing in education.
Zac Boring 2 months ago Analysis
Why You Can't Use Your Right to Try
via LessWrong AI [4] — The Availability Problem:Imagine you have cancer, or chronic pain, or a progressive degenerative disease of some sort. You have exhausted the traditional treatment options available to you, and none of them have worked. However, there are treatments that…
Zac Boring 2 months ago Analysis
Is ProgramBench Impossible?
via LessWrong AI [3] — ProgramBench is a new coding benchmark that all frontier models spectacularly fail. We’ve been on a quest for “hard benchmarks” for a while so it’s refreshing to see a benchmark where top models do badly. Unfortunately, ProgramBench has one big problem:…
Zac Boring 2 months ago Analysis
Claude Code, Codex and Agentic Coding #8
via Substack Zvi [999] — When I started this series, everyone was going crazy for coding agents.
Zac Boring 2 months ago Analysis
The AI industry is where banking was in 2006. (We're hiring)
via LessWrong AI [8] — TL;DR; CeSIA, the French Center for AI Safety is recruiting. French not necessary. Apply by 22 May 2026; Paris or remote in Europe/UK.On August 27, 2005, at an annual symposium in Jackson Hole, Raghuram Rajan, then chief economist of the International…
Zac Boring 2 months ago Analysis
AI #167: The Prior Restraint Era Begins
via Substack Zvi [999] — The era of training frontier models and then releasing them whenever you wanted?
Zac Boring 2 months ago Analysis
Many individual CEVs are probably quite bad
via LessWrong AI [4] — I was thinking about Habryka's article on Putin's CEV, but I am posting my response here, because the original article is already 3 weeks old.I am not sure how exactly a person's CEV is defined. "If we knew everything and could self-modify" seems…
Live Doom Meter
-- %
0% — We're fine 100% — GG
P(Doom) Scoreboard
0%25%50%75%100%
Loading estimates...