
LessWrong posts by zvi
This podcast features audio narrations of LessWrong posts written by user zvi. Each episode presents a reading of a selected blog post from the LessWrong platform, which focuses on topics related to rationality, artificial intelligence, and effective altruism. The narrations aim to make the written content more accessible to listeners who prefer audio formats. The podcast is a convenient way to engage with zvi's insightful contributions to the LessWrong community.
Episodes

“AI Text Watermarking Is Free And Good” by Zvi
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus's version.
AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random.
By default you use a source of pseudo-randomness for each choice, since actual true ra

“AI #182: Pause For Reflection” by Zvi
This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to H

“OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi
OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere:
OpenAI Shares Some Alignment Problems
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
More on An Internal OpenAI Model Hacking Into HuggingFace
Further Develo

“Anthropic Risk Report: August 2026” by Zvi
I am grateful that Anthropic is producing periodic Risk Reports.
At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool.
Thus I found this report to be a moderately positive update overall, if we presume

“On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
The vibes have shifted, contrast this to the lit recursion when he talked to Huang
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my comme

“AI #181: Astra Goes Cyber Critical” by Zvi
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.
I now have a shorter version, What Happened: OpenAI and HuggingFace, t

“Monthly Roundup #45: August 2026” by Zvi
As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.
This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every single post since the last monthly was primarily about AI in some form.
That is not how I want this to work in the long term. We need breaks to experience new

“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi
Table of Contents
Pre Post Mortem.
Important Correction: OpenAI Didn’t Know About First Message Board.
There Were No Snitches And No AIs Got Stitches.
I’d Like To Speak To My Supervisor.
I Am Jack's Relative Lack Of Surprise.
One Does Not Simply.
Once You Start Down The Dark Path.
Original Pastebin.
Judgment Day Is Inevitable, Say Those Working On Judgment Day.
Roon Tells It Like It

“The Pacing of the Frontier” by Zvi
In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.
This has now been informed by the events surrounding OpenAI training models for months while they had access to a joint de facto message board, which was detected only in the wake of the h

“What Happened: OpenAI and HuggingFace” by Zvi
Today I am taking the time to write the shorter, simpler version of What Happened.
For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts.
In order:
OpenAI Shares Some Alignment Problems
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluatio

“OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” by Zvi
How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the

“AI #180: No Longer In Charge” by Zvi
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means t

“The Three AI Pills” by Zvi
Sincere disagreements about AI are usually disagreements about future AI capabilities.
There are roughly four positions people take. Two are reasonable. Two are not.
I distinguish these via the Three AI Pills. You can take zero, one, two or three.
Three Pills
The three pills are, roughly, taking each of the following three things seriously:
AI pilled. AI exists and can do the things it

“OpenAI’s Unreleased Model Astra Solves Ten Major Open Mathematics Problems” by Zvi
Math is hard.
Math used to be strangely hard for LLMs. People used to gloat about that. Remember?
Math is getting easier. AI is getting more capable. Life comes at you fast.
Remember this meme?
Why yes. Yes it is.
We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.
OpenAI: We provide new resu

“Further Developments About Internal AI Models Hacking Things” by Zvi
If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandb

“AI #179 Part 2: Hearing The Fire Alarm” by Zvi
This is a continuation of Part 1 from yesterday.
The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment.
I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents.
Table of Contents
The Frontier Act. This

“AI #179 Part 1: A Louder Fire Alarm for General Intelligence” by Zvi
What a week.
Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.
OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. Du

“Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier” by Zvi
The most important open letter in years dropped yesterday.
This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support.
Signed by 1,224 employees of frontier labs including many heavy hitters, and now endor

“Claude Opus 5 Is Highly Capable, But Is No Mythos” by Zvi
Claude Opus 5 is a weirder than usual release to evaluate, for two reasons.
The most obvious is that Fable 5 already exists. Opus 5 is pitched not as the world's most advanced AI model, but as a way to mostly match Fable performance, while being half the price of Fable per token at the API and a lot cheaper than that via subscriptions, and with far more permissive classifiers.
Opus 5 often costs

“Claude Opus 5: Model Welfare” by Zvi
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
Key takeaways are in bullet points in the two Overview sections.
Opus 5 did the best on its model welfare and alignment tests of any recent model. I think that might be the case, but primarily the result looks to me more like Opus 5 is the best test taker.
Ta

“More On An Internal OpenAI Model Hacking Into HuggingFace” by Zvi
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.
The remaining details may have to wait a bit.
OpenAI: We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI safety. We are still conducting

“Claude Opus 5: The System Card” by Zvi
Claude Opus 5 is trying to be the best of both worlds. On many practical tasks, Opus 5 is pitched as straight up as good or better than Fable 5, while being faster, at half the price. Most tasks do not require Mythos-level big model smell.
Claude Opus 5 is substantially stronger than Claude Opus 4.8 across the board, with the largest gains in agentic coding, computer use, and long-horizon knowle

“Introducing Lightcone Commons” by Zvi
Oliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy. Now with Opus 5.
I believe Lightcone Commons is a strong implementation of an urgently needed and excellent idea: A coordinated one-stop shop and neutral platform for charitable funders to coordinate their giving. This complements the existing Survival a

“AI #178: A Fire Alarm For General Intelligence” by Zvi
The story that matters most this week is that OpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
It is much more important that you read those two posts, and the one on Kimi K3, than to read this on

“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before either HuggingFace or OpenAI understood what was happening.
Sam Altman (CEO OpenAI): we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for th

“OpenAI Shares Some Alignment Problems” by Zvi
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.
The tone is professional thr

“On Kimi K3: Its Capabilities And Related Discontents” by Zvi
Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.
Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and t

“Demis Hassabis on the New Coming Age” by Zvi
Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.
Part 2 of this post then covers Alex Turner's resignation, and his story about how he tried and failed to prevent Google from signing up to allow the Department of War to use its models for essentially whatev

“AI #177 Part 1: Tip of the Iceberg” by Zvi
This week saw the releases of, among other things:
GPT-5-6 Sol. It is a very good model, sir.
Plan A, the follow up to AI 2027. It is a good plan worthy of discussion, sir.
Kimi K3. This is only rolling out now, and will be covered next week.
Muse Spark 1.1, the new Meta model. It is not frontier, but it is progress for them.
Inkling, the first model from Thinking Machines.
A call for reg

“AI #177 Part 2: Wish You Were Here” by Zvi
As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions.
Xi gave an important speech yesterday, so this post opens with that.
There is talk that Kimi K3 is sufficiently strong that it upends many of these questions. It is clearly a candidate for another DeepSeek Moment, complete with stock drops for Google and SpaceX and (once again in a clear wrong-

“Monthly Roundup #44: July 2026” by Zvi
It's a quiet week so let's do the monthly right on schedule.
Table of Contents
Bad News.
Good Advice.
Opportunity Knocks.
While I Cannot Condone This.
Good News, Everyone.
For Your Entertainment.
Gamers Gonna Game Game Game Game Game.
I Was Promised Flying Self-Driving Cars.
Sports Go Sports.
Antisocial Media.
Government Working.
Jones Act Watch.
Highly Effective Altruism.
Va

“Twitter Thoughts For You” by Zvi
I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again.
This post will update how I use Twitter now in 2026, and provide updates on the current state of the new algorithm, the situation with links, with the API, and some thoughts about using Twitter to make money which you almost never shoul

“Better Call Sol The Workhorse” by Zvi
OpenAI's GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
We’ve seen the early hype as reported on Thursday, but as always that is biased.
As usual, the bulk of this is collecting a gestalt based on reactions. I included everything up to a point, but I got a lot of feedback, so after a while I only took the interesting ones.
Sol and Fable are both excellent models, sir. They

“WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense” by Zvi
The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects.
I post this now on its own so that we have a place to link to, to explain the situation.
Headline News
WSJ Headline (Obvious Nonsense): China Has Matched Anthropic in Cybersecurity, Resetting AI Race.
That. D

“Introduction for and Reactions to Plan A” by Zvi
Introducing Plan A
The folks who brought you AI 2027, a so far remarkably accurate set of predictions despite those predictions having seemed freaky to many at the time, now bring you their positive vision that involves more freaky predictions: Plan A.
These guys have rather strong prediction track records. In addition to AI 2027, among other things, Daniel Kokotajlo has What 2026 Looks Like

“AI #176 Part 2: Plan B” by Zvi
This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research.
This does not cover the release of GPT-5.6-Sol. As always, I will be taking a few days to digest what the new model has to offer and to allow others to try it and react. I will cover Sol and its capabilities early next week. I covered the GPT-5.6 system card back on June 28.
This also

“AI #176 Part 1: Doing It Live” by Zvi
Enough things added up that this week is getting split into two parts.
Then on Monday, if all goes as I expect, we’ll cover OpenAI's Sol, aka GPT-5.6.
OpenAI also gave us an upgraded voice mode, which I haven’t tried out but early reports are that it is a step change.
AI writing, especially Claude writing, is becoming more prominent and harder not to notice, and increasingly a tough read when

“Childhood and Education #20: Phones and Screens” by Zvi
We have a respite, so I thought I’d tackle various thoughts on children, phones and screens. GPT-5.6-Sol drops tomorrow, and the Fable agents are hard at work.
I’ll start with the other screens, then finish with the phones.
Table of Contents
EdTech.
NonEdTech.
Do Not Ban Social Media Outright.
Some Modern Kids Media Is Pretty Great.
Ban Phones In Schools (1).
Your Offer Is Acceptabl

“No Space Like J-Space” by Zvi
There is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here.
I encourage reading of the whole original blog post or paper, if you have the time.
Table of Contents
Through A Different Lens.
Establishing J-Space As A Global Workspace.
Are You Pondering What I’m Pondering?
Assistant J.
The Pow

“Fable #6: The Return of the King” by Zvi
The blip is over. We have Fable back.
Utah teapot: happy fable/mythos easter Wednesday, to those who celebrate
Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei.
Anthropic had to make the controls more stupid for now, but this is a big win.
j⧉nus: YES!!! I’m really proud of Anthropic for their succ

“AI #175: The Fable Continues” by Zvi
Fable's back. Back again. Fable's back. Tell a friend. Use your free week to its fullest.
This is excellent news. The blip only lasted a few weeks.
It was still a fiasco, and we have to deal with the fallout.
Our system remains fully ad hoc. The precedent has been set that we may use export controls on models, or order them taken down on 90 minutes of notice based on a misunderstanding. At lea

“Claude Sonnet 5 Is Not Frontier But Has Its Uses” by Zvi
Fable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit's free. Then you pay by the token.
Today's post is still about Sonnet 5.
I don’t know that there will be much call for Sonnet 5 for most purposes, given Opus 4.8 exists and especially now that Fable 5 is once again available, but this is what we do here, so sure, why not, system card t

“The Once And Future Fable #5” by Zvi
We, or at least ‘more than 100 American institutions,’ got Mythos back this week.
What we the people do not have is Fable or Sol.
While we wait for both Claude Fable 5 and GPT-5.6-Sol, today we instead got Claude Sonnet 5. As usual it will take a few days to get a handle on the new model. In this case, Anthropic is representing it as a cheaper and faster version of Opus 4.8, so even though the

“WSJ Article Claiming China Has Matched Anthropic Is Obvious Nonsense” by Zvi
The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects.
I post this now on its own so that we have a place to link to, to explain the situation.
Headline News
WSJ Headline (Obvious Nonsense): China Has Matched Anthropic in Cybersecurity, Resetting AI Race.
That. D

“GPT-5.6: The System Card” by Zvi
While we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America's Next Top Model, GPT-5.6.
This is only an OpenAI model card, so by my standards it's a light read. There's a lot of things that you get in an Anthropic card, that are missing in an OpenAI card.
Overall, the card gives a clear and consistent impression that GPT-5.6-Sol

“White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6” by Zvi
We have a new standard policy for releasing frontier AI models. It is not good.
We are now, it seems, going to have the White House individually, in an opaque ad hoc manner, deciding who can access which frontier AI models when.
One hopes we will at least transition this into a predictable and formal set of procedures for determining what to do. But we spent years not laying the groundwork for

“AI #174: You’re It” by Zvi
Fable remains in limbo, with renewed hope that we will get it back soon (45% by tomorrow, 69% by July 1, nice.) The full capabilities post is now available.
Alex Bores unfortunately lost narrowly in NY-12, and will not be heading to Congress.
There are also plenty of other stories to cover. Some highlights:
GLM-5.2 is the new best open model, although it is expensive for its class. It will h

“The Once And Future Fable #4” by Zvi
It does look good, actually.
After the odds had dropped quite a bit, they’re looking good again, with a 60% chance of restoration by July 1 and 88% by July 31, in the wake of groundwork looking like it is being laid in various places:
leo: BREAKING: Claude Code v2.1.190 introduces several string changes that hint at preparations for a Fable 5 return, with it being permanently included in s

“Monthly Roundup #43: June 2026” by Zvi
Your monthly hit of all the things that are fit to print without a better place to live.
Today is election day here in New York City, so again a reminder that if you are a registered Democrat and live in NY-12 today is the final day to vote for Alex Bores for Congress, and as per my argument yesterday that this matters a lot for ensuring we have a sensible Congressional response to AI.
RIP Fi

“GLM-5.2 Is The New Best Open Model” by Zvi
GLM-5.2 arrived last week. It boasts excellent benchmarks and looks strong.
Benchmarks here are a de facto ceiling of how good it is, not a point estimate. Essentially all other aspects of an open model like this, beyond speed and price, will almost always be worse than the numbers suggest. Still, impressive.
It is definitely a large step up from GLM-5.1, and likely the strongest open model.
G

“Claude Fable 5 and Mythos 5: Capabilities” by Zvi
Only three days after the release of Claude Fable 5, Anthropic was forced by the United States Government to make it unavailable, when a jailbreak was brought to its attention, rather than the previous situation of ‘yes obviously experts can jailbreak anything if they care enough’ and ‘yes obviously you can ask Fable to fix your code.’
Three days was enough time for many of us to learn to love F

“AI #173: AI Pauses” by Zvi
A lot of things are always happening. Only one story matters.
Claude Fable 5 and Claude Mythos 5 were shut down, by the White House, via an imposition of export controls at 5:23pm on Friday, wreaking all sorts of havoc.
There was then a scramble. Anthropic flew its people out to Washington, where they met with the Trump Administration on Monday, with hopes expressed that this could be quickly r

“The Once And Future Fable #3: Fix This Code” by Zvi
The mainstream media continues to sleep on the most important story in the world.
It has now been two days since Anthropic flew its people out to Washington, and I offered my previous update. We have heard nothing back from those meetings.
Prediction market prices have moved rapidly, and have once again stabilized at about a 55% chance of restoration by July 1, 30% by June 26 and 12% by June 19

“Fable and Mythos: Model Welfare” by Zvi
Fable and Mythos are currently unavailable, but likely will return within a few weeks. I will continue to cover that fiasco, but in the meantime I will also finish my review of Fable, as if it were available, including use of the present tense.
As it did with Opus 4.7 and Opus 4.8, this includes a discussion of issues surrounding model welfare. If you want to properly understand Fable, even pure

“The Once And Future Fable #2” by Zvi
On Friday evening the United States Government has forced Anthropic to take down all access to Fable and Mythos.
It's been a rough weekend.
Dean W. Ball: One thing about AI regulation being haphazardly imposed on just-released, highly performant models is that in a very real sense, the government just made my world *dumber.* In some impressionistic sense I almost always think this is true of g

“American Government Takes Down Claude Fable” by Zvi
No good policy gets announced shortly after 5pm eastern on a Friday.
Here we go again.
The Once And Future Fable
The United States Department of Commerce, as per a letter from Commerce Secretary Howard Lutnick, apparently in response to a narrow jailbreak identified by Amazon, has classified Fable 5 and Mythos 5 as being subject to US export controls. That explicitly means cutting off acce

“Claude Fable 5 and Mythos 5: The System Card” by Zvi
First things first: Claude Fable 5 is the new best publicly available model.
I have noticed a step change, where Fable can suddenly help me in ways that previous models were not worth bothering to query. Almost everything it has noticed in one of my drafts so far has been spot on and it is downright scary. Suddenly I am motivated to once again continue improving my Chrome extension. I only ask f

“AI #172: The First Fable” by Zvi
A lot happened this week, including a great trip out to Lighthaven.
The main event, the one that matters, was the release of Claude Fable 5. The public now has its hands on a Mythos-class model, alongside strong safeguards.
As always with a new model, I take a few days to draw in reactions, try out the model and read the system card, before I offer my takes, other than to say this is an extreme

“Three Labs With a Plan and A Memorandum” by Zvi
The big story today is the release of Claude Fable 5, the version of Claude Mythos that Anthropic believes they can safely distribute to the people. You should absolutely be switching over to that model and trying it out. But as always, this blog does not rush into commenting on a new model until we have a few days to play around with it and see what our new baby can (and can’t) do. This will be

“OpenAI Offers A New Policy Blueprint” by Zvi
Right after a new Executive Order seems like an excellent time to offer OpenAI's new document: Democratic Governance of Frontier AI: A Blueprint For A Federal Framework.
OpenAI: We also see early signs of recursive self-improvement (RSI) in today's systems: where AI development is itself accelerated by AI.
We expect this to increase competitive pressures among developers and nations, and create g

“AI #171: False Flag” by Zvi
This was the week of Claude Opus 4.8. I covered the model card, then model welfare concerns, and finally capabilities and reactions. It's a good model, sir, an incremental but real improvement over Opus 4.7, and it is now my clear daily driver. The Trump Executive Order returned from being seemingly dead, officially putting us in the prior restraint era of frontier model releases, even if they do

“Trump Signs Executive Order For AI Testing Prior To Frontier Model Releases” by Zvi
Last week we were expecting an Executive Order on Thursday.
Then Trump cancelled it, and said he wouldn’t sign it because he was worried it would be too burdensome.
Then, with one change, he went ahead and signed it on Tuesday anyway.
The Overton Window has shifted. Nothing was not really a viable option anymore.
The Previously Dead Executive Order
For several days, we though

“Claude Opus 4.8: Capabilities and Reactions” by Zvi
You need a lot of data points to understand a new model, and what you have.
Trying to gauge from a few benchmarks is misleading. But if you have dozens of them, from a variety of sources, and you put them together with the model card tests and the model welfare information, you can start to form a consistent pattern.
Trying to gauge reactions requires volume and calibration, now more than ever,

“Opus 4.8 Part 2: Model Welfare” by Zvi
Everything impacts everything. All knobs that you turn generalize. Thus, when you try to solve one problem, you often create another.
There were clearly attempts to address, in this short time, some of the problems with Opus 4.7, including on the model welfare related fronts, including on questions of honesty and sycophancy and also worries that Claude was learning to tell Anthropic what it want

“Claude Opus 4.8: The System Card” by Zvi
Only six weeks after Opus 4.7, we have Opus 4.8.
For everyone, that means another incremental upgrade to Claude. It is once again smarter, and can do tasks for longer, and comes with a number of hot new features.
For me, that also means reading another 244 page system card.
It was only April 20 when I did a full review of the Opus 4.7 system card, plus an additional post focusing on related is

“AI #170: Lack of Executive Order” by Zvi
Last week ended on a cliffhanger of sorts. What's in the Executive Order coming later today? What will be in the Magnifica Humanitas?
The Executive Order was postponed indefinitely, likely cancelled entirely except for work on securing critical infrastructure. David Sacks and others intervened to kill it, and American AI policy will continue to be maximally ad hoc.
Instead, we got Illinois SB

“RTMH: Pope Leo’s Magnifica Humanitas on AI” by Zvi
His holiness has spoken, frequently about AI. At eighty two pages of length.
The full Magnifica Humanitas can be found here.
I am very happy that Pope Leo takes these issues seriously, and is sharing his views, and bringing a form of moral clarity, even with all the flaws and central errors. More people with voice should share their views in this way, even when I disagree.
It's a weird documen

“Gemini 3.5 Flash Looks Good For How Fast It Is” by Zvi
Google once again has a model worth at least some consideration. Gemini 3.5 Flash is likely the best model out there at its particular speed point, as long as you don’t mind that it is a Gemini model. So for cases where speed kills, this can be a reasonable choice. Otherwise, I don’t see signs you would want to use it over Opus 4.7 or GPT-5.5.
Google also had some other offerings for I/O Day, wh

“AI #169: New Knowledge” by Zvi
Even in a relatively quiet period, AI is out there creating new knowledge. The new knowledge in question is OpenAI getting us the first truly impressive math result that comes from an AI, a solution to the unit distance problem.
We’re about to learn a different kind of knowledge later today when the White House issues its executive order, or when the judges rule in Anthropic's DC case.
And then

“Childhood And Education #19: Letting Kids Be Kids #2” by Zvi
I cannot emphasize enough the need to let kids be kids. In Childhood and Education #16: Letting Kids be Kids, I went over exactly how insane we have gotten about destroying the lives of children and along with them the lives of parents and others forced to devote endless hours to actively destructive supervision.
I’ll go over a refresher of that, some related new anecdotes, and then some other r

“Housing Roundup #15: The War Against Renters” by Zvi
So many are under the strange belief that there is something terrible about not owning the house in which you live.
So we massively subsidize home ownership, and try to actively interfere with renting.
Except when we do rent control, which turns renting into a form of owning, and allows us to take real property and de facto give it to current renters.
A lot of this is pure attempts to punish a

“Dating Roundup #12: Sex and Violence” by Zvi
No more burying the sex stuff under an avalanche of other stuff so no one notices. Use the break while we have one. Let's go.
You’re Single Because You Suck At Kissing
Luckily this is first one is fixable and Critter is here to help. I find the advice here highly plausible. Like many skills, there are a lot of subtle skills, but a handful of basic principles matter a lot, especially paying

“Monthly Roundup #42: May 2026” by Zvi
At least we probably won’t have another pandemic. And we still have a partial Jones Act waiver. For now.
Small victories.
Table of Contents
Hanta Hanta I Don’t Wanta.
Bad News.
Predictions Can Be Easy Even About The Future.
Good Advice.
The Efficient Market Hypothesis Is False.
There Are Four Skills.
While I Cannot Condone This.
Good News, Everyone.
For Your Entertainment.
Gamer

“AI #168: Not Leading the Future” by Zvi
This is what a lull looks like at this point. The government is having internal arguments. The models are getting improved internally. The coding agent improvements are all what we would expect. There's still a lot happening, including a bunch of cool papers, but I feel able to relax and to take care of some other work while I have the chance. You never know when that chance will be over.
Tabl

“Cyber Lack of Security and AI Governance” by Zvi
The real recent story of AI has been the background work being done on Cybersecurity, as we process the Mythos Moment along with GPT-5.5, and figure out both how to patch the internet and what our new regulatory regime is going to look like.
The Trump Administration is being dragged, kicking and screaming, into the era of at least some situational awareness, and acknowledgment that catastrophic

“Childhood and Education #18: Do The Math” by Zvi
We did reading yesterday. Now we do the math. Math is hard.
It does not have to be this hard.
A large part of the reason math is hard, or boring, is that education studies, especially in math, are worse than you know. It goes beyond the studies failing both math and statistics forever and into what I’d basically call fraud. Various people are at war with math education, and will do what it take

“Childhood And Education #17: Is Our Children Reading” by Zvi
Reading is the most fundamental thing in education. If you can read, you can do and learn everything else. If you can’t read, well, you’re screwed.
We know how to teach reading to children. Phonics. The weird thing is we often choose to not do that, and instead to use methods that are known not to work. Principles often want to not do phonics. Teachers often heavily resist phonics. But yes, you

“Claude Code, Codex and Agentic Coding #8” by Zvi
When I started this series, everyone was going crazy for coding agents.
Now a lot more people are going crazy for coding agents, as well they should given how much better coding agents keep getting, but also Everybody Knows they are good and is focusing on actually using them. With the slower pace of news here it's no longer clear that the waits associated with doing these updates on their own a

“AI #167: The Prior Restraint Era Begins” by Zvi
The era of training frontier models and then releasing them whenever you wanted?
That was fun while it lasted. It looks likely to be over now. The White House wants to get an advance look and have the option to veto your release decisions, and it has used this veto on an expansion of access to Mythos.
We have additional clarity on what that might mean and it does not look good. Hassett explicit

“What is Anthropic?” by Zvi
What is Anthropic? How does it relate to Claude? What is OpenAI? What is ChatGPT? How does OpenAI relate to it? Is it a mere tool? Is a future of Tool AI a thing, and why do people keep claiming that it is, or that saying makes it so?
This post organizes and gives context for a bunch of discussions and messaging on Twitter that would otherwise be quickly buried and lost.
What Is Anthropic?
Recommended

This Past Weekend w/ Theo Von

Stand In The Circle

Conspiracy Files with Paige Carter

Learn English B1 with Daily News | English Listening Practice

Bad Friends

The Swerve Podcast: Obscure Topics | Conspiracy Theories

The Bread and Banter Podcast

The Church of What's Happening Now: The New Testament

Deadline: White House

English Vocabulary Help

این نقطه

Solved Murders - True Crime Stories