Home Podcasts LessWrong posts by zvi
LessWrong posts by zvi

LessWrong posts by zvi

zvi 250 Episodes Aug 21, 2026

This podcast features audio narrations of LessWrong posts written by user zvi. Each episode presents a reading of a selected blog post from the LessWrong platform, which focuses on topics related to rationality, artificial intelligence, and effective altruism. The narrations aim to make the written content more accessible to listeners who prefer audio formats. The podcast is a convenient way to engage with zvi's insightful contributions to the LessWrong community.

Episodes

“AI Text Watermarking Is Free And Good” by Zvi
“AI Text Watermarking Is Free And Good” by Zvi Aug 21, 2026 1415 Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner. Here is how his solution works, or see Tenobrus's version. AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random. By default you use a source of pseudo-randomness for each choice, since actual true ra
“AI #182: Pause For Reflection” by Zvi
“AI #182: Pause For Reflection” by Zvi Aug 20, 2026 5259 This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward. OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to H
“OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi
“OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi Aug 19, 2026 2346 OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Develo
“Anthropic Risk Report: August 2026” by Zvi
“Anthropic Risk Report: August 2026” by Zvi Aug 18, 2026 4744 I am grateful that Anthropic is producing periodic Risk Reports. At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool. Thus I found this report to be a moderately positive update overall, if we presume
“On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi
“On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi Aug 15, 2026 2956 Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. The vibes have shifted, contrast this to the lit recursion when he talked to Huang As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my comme
“AI #181: Astra Goes Cyber Critical” by Zvi
“AI #181: Astra Goes Cyber Critical” by Zvi Aug 13, 2026 6099 The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew. I now have a shorter version, What Happened: OpenAI and HuggingFace, t
“Monthly Roundup #45: August 2026” by Zvi
“Monthly Roundup #45: August 2026” by Zvi Aug 12, 2026 2211 As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every single post since the last monthly was primarily about AI in some form. That is not how I want this to work in the long term. We need breaks to experience new
“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi
“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi Aug 11, 2026 3251 Table of Contents Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My Supervisor. I Am Jack's Relative Lack Of Surprise. One Does Not Simply. Once You Start Down The Dark Path. Original Pastebin. Judgment Day Is Inevitable, Say Those Working On Judgment Day. Roon Tells It Like It
“The Pacing of the Frontier” by Zvi
“The Pacing of the Frontier” by Zvi Aug 10, 2026 2138 In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so. This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the h
“What Happened: OpenAI and HuggingFace” by Zvi
“What Happened: OpenAI and HuggingFace” by Zvi Aug 8, 2026 1247 Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts. In order: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluatio
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi Aug 7, 2026 4735 How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know? At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the
“AI #180: No Longer In Charge” by Zvi
“AI #180: No Longer In Charge” by Zvi Aug 6, 2026 4204 What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means t

Recommended