Home Podcasts AXRP - the AI X-risk Research Podcast
AXRP - the AI X-risk Research Podcast

AXRP - the AI X-risk Research Podcast

Daniel Filan 63 Episodes Aug 2, 2026

AXRP (pronounced axe-urp) is the AI X-risk Research Podcast hosted by Daniel Filan. In each episode, Daniel talks with researchers about their papers on reducing the risk of artificial intelligence causing an existential catastrophe. The conversations explore why the research was done and how it might help prevent AI from permanently and drastically curtailing humanity's future potential. Transcripts and more information are available at axrp.net.

Episodes

50 - Eli Lifland on AI 2027
50 - Eli Lifland on AI 2027 Aug 2, 2026 02:40:40 Remember AI 2027? Not AI 2040, the newest coolest thing AI Futures Project has done, but AI 2027, their OG product? At long last, we have an AXRP episode about it. Enjoy!   Transcript: https://axrp.net/episode/2026/08/03/episode-50-eli-lifland-ai-2027.html   Topics we discuss, and timestamps: 0:00:12 What is AI 2027? 0:10:17 What happens in AI 2027? 0:18:08 Why two endings? 0:21:15 Who did what?
49 - Caspar Oesterheld on Program Equilibrium
49 - Caspar Oesterheld on Program Equilibrium Feb 18, 2026 02:32:07 How does game theory work when everyone is a computer program who can read everyone else's source code? This is the problem of 'program equilibria'. In this episode, I talk with Caspar Oesterheld on work he's done on equilibria of programs that simulate each other, and how robust these equilibria are. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/axrpodcast Transcript: http
48 - Guive Assadi on AI Property Rights
48 - Guive Assadi on AI Property Rights Feb 15, 2026 02:05:34 In this episode, Guive Assadi argues that we should give AIs property rights, so that they are integrated in our system of property and come to rely on it. The claim is that this means that AIs would not kill or steal from humans, because that would undermine the whole property system, which would be extremely valuable to them. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/a
47 - David Rein on METR Time Horizons
47 - David Rein on METR Time Horizons Jan 2, 2026 01:47:17 When METR says something like "Claude Opus 4.5 has a 50% time horizon of 4 hours and 50 minutes", what does that mean? In this episode David Rein, METR researcher and co-author of the paper "Measuring AI ability to complete long tasks", talks about METR's work on measuring time horizons, the methodology behind those numbers, and what work remains to be done in this domain. Patreon: https://www.pat
46 - Tom Davidson on AI-enabled Coups
46 - Tom Davidson on AI-enabled Coups Aug 7, 2025 02:05:26 Could AI enable a small group to gain power over a large country, and lock in their power permanently? Often, people worried about catastrophic risks from AI have been concerned with misalignment risks. In this episode, Tom Davidson talks about a risk that could be comparably important: that of AI-enabled coups. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/axrpodcast Transc
45 - Samuel Albanie on DeepMind's AGI Safety Approach
45 - Samuel Albanie on DeepMind's AGI Safety Approach Jul 6, 2025 01:15:42 In this episode, I chat with Samuel Albanie about the Google DeepMind paper he co-authored called "An Approach to Technical AGI Safety and Security". It covers the assumptions made by the approach, as well as the types of mitigations it outlines. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/axrpodcast Transcript: https://axrp.net/episode/2025/07/06/episode-45-samuel-albani
44 - Peter Salib on AI Rights for Human Safety
44 - Peter Salib on AI Rights for Human Safety Jun 28, 2025 03:21:33 In this episode, I talk with Peter Salib about his paper "AI Rights for Human Safety", arguing that giving AIs the right to contract, hold property, and sue people will reduce the risk of their trying to attack humanity and take over. He also tells me how law reviews work, in the face of my incredulity. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/axrpodcast Transcript: ht
43 - David Lindner on Myopic Optimization with Non-myopic Approval
43 - David Lindner on Myopic Optimization with Non-myopic Approval Jun 15, 2025 01:40:59 In this episode, I talk with David Lindner about Myopic Optimization with Non-myopic Approval, or MONA, which attempts to address (multi-step) reward hacking by myopically optimizing actions against a human's sense of whether those actions are generally good. Does this work? Can we get smarter-than-human AI this way? How does this compare to approaches like conservativism? Listen to find out. Patr
42 - Owain Evans on LLM Psychology
42 - Owain Evans on LLM Psychology Jun 6, 2025 02:14:26 Earlier this year, the paper "Emergent Misalignment" made the rounds on AI x-risk social media for seemingly showing LLMs generalizing from 'misaligned' training data of insecure code to acting comically evil in response to innocuous questions. In this episode, I chat with one of the authors of that paper, Owain Evans, about that research as well as other work he's done to understand the psycholog
41 - Lee Sharkey on Attribution-based Parameter Decomposition
41 - Lee Sharkey on Attribution-based Parameter Decomposition Jun 3, 2025 02:16:11 What's the next step forward in interpretability? In this episode, I chat with Lee Sharkey about his proposal for detecting computational mechanisms within neural networks: Attribution-based Parameter Decomposition, or APD for short. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com/axrpodcast Transcript: https://axrp.net/episode/2025/06/03/episode-41-lee-sharkey-attribution-ba
40 - Jason Gross on Compact Proofs and Interpretability
40 - Jason Gross on Compact Proofs and Interpretability Mar 28, 2025 02:36:05 How do we figure out whether interpretability is doing its job? One way is to see if it helps us prove things about models that we care about knowing. In this episode, I speak with Jason Gross about his agenda to benchmark interpretability in this way, and his exploration of the intersection of proofs and modern machine learning. Patreon: https://www.patreon.com/axrpodcast Ko-fi: https://ko-fi.com
38.8 - David Duvenaud on Sabotage Evaluations and the Post-AGI Future
38.8 - David Duvenaud on Sabotage Evaluations and the Post-AGI Future Mar 1, 2025 20:42 In this episode, I chat with David Duvenaud about two topics he's been thinking about: firstly, a paper he wrote about evaluating whether or not frontier models can sabotage human decision-making or monitoring of the same models; and secondly, the difficult situation humans find themselves in in a post-AGI future, even if AI is aligned with human intentions.   Patreon: https://www.patreon.com/axrp

Recommended