
Inside the Black Box: Cracking AI and Deep Learning
Inside the Black Box: Cracking AI and Deep Learning explores how large language models like ChatGPT actually work. It breaks down core ideas in artificial intelligence, neural networks, and deep learning in an accessible way. The show is hosted by Arshavir Blackwell, PhD.
Episodes

Not at This Address
This episode challenges the old idea of a single language box in the brain, using aphasia studies and hospital control groups to show how language processing is more distributed than once thought. It also explores how different languages rely on word order or verb marking in different ways, and why those differences matter when the brain is under stress.

My Advisors Argued This for Thirty Years. Now You Can Check
Thirty years ago, the authors of Rethinking Innateness argued that grammar could come out of a learner with no grammar built in. Nobody could check it. Now there's an instrument. A reading of the Inside the Black Box essay on what Elman, Bates, and Karmiloff-Smith would make of transformer language models — attention as a lookup, structure nobody installed, and the part that isn't there.
Read the

Inside the Past Tense: Raw Letters, Raw Sound
Can a model learn past tense the way children seem to, if you remove the tokenizer from the equation? This final installment tests raw letters and raw sound, probes whether a U-shaped pattern can be forced, and shows why the answer points to distributed, frequency-driven learning rather than a discrete rule.

Forcing a Neural Network to Add -ed
Part 3 of 4. After two episodes arguing there is no discrete past-tense rule inside a language model, this one turns the strongest possible search on the question: gradient descent, hunting for the single internal direction that best forces goed over went.
It works — 100% of the time, even on held-out verbs it was never tuned on. Then one control collapses the whole result. That same direction dr

The Wug Test for AI
This episode explores a modern twist on Jean Berko’s famous wug test, comparing how large language models handle made-up verbs versus familiar ones. It digs into tokenization, internal feature probes, and why neural networks seem to approximate grammar through probabilistic patterns rather than clean symbolic rules.

Learning the Past Tense in AI
This episode revisits the famous debate over whether language is learned through symbolic rules or distributed neural patterns, using the classic irregular verb U-shaped curve as the battleground. It then compares child language development with training snapshots from modern AI models, showing that transformers learn past tense in a strikingly different way: starting with the broad rule and gradu

Fine Tuning Lora: It's Not What You Think
When you fine-tune an AI model, what changes inside doesn't predict what changes outside. This week on Inside the Black Box, I break down why — and what it means for anyone auditing or regulating these systems.

When Fluent Answers Start Sounding True
This episode explores why smooth, coherent language can feel more credible than it is, and how processing fluency, familiarity, and authority cues shape what we believe. It also digs into why conversational AI is especially persuasive, from polished explanations to confident-sounding confabulations.

Why Your Brain Believes the Model
The Heuristic Loop You Can't Break from Inside

When Polished Answers Feel Finished
This episode explores fluency-as-validity: the way polished AI responses can make us feel like the work of judgment is already done. It also looks at why large language models are so effective at creating the sensation of clarity, and why mechanistic interpretability may be a way to push back against that enchantment.

What Seneca Teaches Us that Marcus Couldn't
716 features fire on both Seneca and Marcus Aurelius but stay dark for ad copy. The model learned Stoic philosophy, not just an author's style. Plus: why 'inert' features aren't all the same thing.

The Pattern Holds for Another Author
We trained a fresh LoRA on the letters of Seneca and ran the same analysis pipeline we used on Marcus Aurelius and advertising copy. Every structural finding replicated. The model organizes its adaptation into five clusters: one tight (features moving in lockstep) and four loose (features cooperating more independently). Seneca produced the cleanest clustering we've measured and the strongest work

The Pattern Holds
We replicated our Marcus Aurelius findings at a new layer, then threw the whole method at 12 commercial ad copy styles trained into a single LoRA. The patterns held, and the new domain revealed something we couldn't have seen before: the model organizes its adaptations by register family, not by individual style.

Cracking Open the Black Box
We opened the 65%. The features that resisted interpretation one at a time turned out to organize into five co-activation clusters with clear thematic identities and causal effects nearly ten times stronger than any individual feature. Second in a series with John Holman.

Inside a Fine-Tuned Language Model
A concise, single-segment episode of Inside the Black Box: Cracking AI and Deep Learning where Arshavir Blackwell explains, in one continuous narrative, what neural networks are, how their simple units combine into powerful systems, and how learning by backpropagation sculpts their behavior. This short episode is designed as an elegant, one-paragraph-style monologue that introduces listeners to ne

What Counts as Structure? From Harris and Elman to Today’s Neural Nets
This episode of Inside the Black Box: Cracking AI and Deep Learning tells the story of an unexpected convergence in the history of language and AI. In 1995, Peter Bensch noticed that Zelig Harris, a mid‑century structural linguist, and Jeff Elman, a pioneer of simple recurrent networks, had independently uncovered the same deep insight about language: structure lives in patterns of use.
Arshavir

Building a House Without Blueprints: When Interpretability Tools Work — and When They Don’t
This episode of Inside the Black Box: Cracking AI and Deep Learning explores a new theoretical framework that unifies sparse autoencoders (SAEs), transcoders, and crosscoders — and what it tells us about when mechanistic interpretability actually works.
We start by demystifying these tools and how they use sparse features to uncover internal concepts and computations in large language models, fro

I Told My LLM Not to Say "Empower"
In this episode of Inside the Black Box: Cracking AI and Deep Learning, Arshavir Blackwell, PhD, takes engineers and researchers inside the practical mechanics of LoRA, low‑rank adaptation methods that make it possible to fine‑tune multi‑billion‑parameter language models on a single GPU.

Beyond the Surface of AI Intelligence
This episode dives into why judging AI by behavior alone falls short of proving true intelligence. We explore how insights from mechanistic interpretability and cognitive science reveal what’s really happening inside AI models. Join us as we challenge the limits of behavioral tests and rethink what intelligence means for future AI.

Unlocking BERTs Hidden Grammar
Explore how BERT’s attention heads reveal an emergent understanding of language structure without explicit supervision. Discover the role of attention as a form of memory and what it means for the future of AI language models.

Cracking the Code of AI Interpretation
Dive into how we naturally explain neural networks with folk interpretability and why these simple stories fall short. Discover the journey toward mechanistic understandability in AI and what that means for how we talk about and trust large language models.

Decoding GPTs Hidden Circuits
Explore how sparse autoencoders and transcoders unveil the inner workings of GPT-2 by revealing functional features and computational circuits. Discover breakthrough methods that shift from observing raw network activations to mapping the model's actual computation, making AI behavior more interpretable than ever.

Decoding Attention and Emergence in AI
Explore how attention heads uncover patterns through learned queries and keys, revealing emergent behaviors shaped by optimization. Dive into parallels with natural selection and psycholinguistics to understand how meaning arises not by design but through experience in both machines and brains.

When Knowledge Battles Noise in GPT Models
Explore how GPT-2 balances fleeting factual recall with generic responses through internal competition among candidate answers. Discover parallels with human cognition and how larger models navigate indirect recall to reveal hidden knowledge beneath suppression.

Inside Circuits: How Large Language Models Understand
Dive into the world of neural circuits within large language models. In this episode, Arshavir Blackwell unpacks how transformer circuits, attention mechanisms, and high-dimensional geometry combine to create the magic—and limits—of modern AI language systems.

Hallucinations, Interpretability, and the Seahorse Mirage
This episode dives into why advanced language models still generate hallucinations, how interpretability tools help us uncover their hidden workings, and what the seahorse emoji teaches us about model and human reasoning. Arshavir connects groundbreaking research, practical business importance, and the statistical quirks that shape AI's version of 'truth.'

How Transformers Stack Meaning Like Finnish Words
Explore how large language models build up meaning in ways strikingly similar to the layered grammar of Finnish. Arshavir Blackwell reveals why understanding Finnish morphology offers a powerful analogy for interpreting the compositional logic inside modern AI systems.

The Mandela Effect in AI: Why Language Models Misremember
Dive into how and why large language models like ChatGPT mirror the human Mandela Effect, reproducing our collective false memories and misquotations. Arshavir Blackwell examines the science behind errors in models and minds, and explores how new techniques can counteract these uncanny AI confabulations.

Bridging Circuits and Concepts in Large Language Models
How do millions of computations inside large language models add up to something like understanding? This episode explores the latest breakthroughs in mechanistic interpretability, showing how tools like representational geometry, circuit decomposition, and compression theory illuminate the missing middle between circuits and meaning. Join Arshavir Blackwell as he opens the black box and challenge

How Transformers Turn Words Into Meaning
Embark on a step-by-step journey through the inner workings of transformer models like those powering ChatGPT. Arshavir Blackwell breaks down how context, attention, and high-dimensional geometry turn isolated tokens into fluent, meaningful language—revealing the mathematics of understanding inside the black box.

Can Smaller Language Models Be Smarter?
Today we explore whether mechanistic interpretability could hold the key to building leaner, more transparent—and perhaps even smarter—large language models. From knowledge distillation and pruning to low-rank adaptation, we examine cutting-edge strategies to make AI models both smaller and more explainable. Join Arshavir as he breaks down the surprising challenges of making models efficient witho

The Weird Geometry That Makes AI Think
Explore how large language models use high-dimensional geometry to produce intelligent behavior. We peer into the mathematical wilderness inside transformers, revealing how intuition fails, and meaning emerges.

Can We Fix It?
Arshavir Blackwell takes you on a journey inside the black box of large language models, showing how cutting-edge methods help researchers identify, understand, and even fix the inner quirks of AI. Through concrete case studies, he demonstrates how interpretability is evolving from an arcane art to a collaborative science—while revealing the daunting puzzles that remain. This episode unpacks the s

Using Symbolic AI to Explain LLMs
Delve into the mysterious world of neural circuits within large language models. We’ll dismantle the jargon, connect these abstract ideas to real examples, and discuss how circuits help bridge the gap between machine learning and human cognition.

Peering Inside the Black Box
Mechanistic interpretability and artificial psycholinguistics are transforming our understanding of large language models. In this episode, Arshavir Blackwell explores how probing neural circuits, behavioral tests, and new tools are unraveling the mysteries of AI reasoning.
Recommended

Fantasy Flex

Solved Murders - True Crime Stories

紐約鳥|New York Aperture

Back to the 80s Radio

The Swerve Podcast: Obscure Topics | Conspiracy Theories

The Bread and Banter Podcast

The Church of What's Happening Now: The New Testament

Crime Stories with Nancy Grace

Deadline: White House

این نقطه

TED Business

Dateline NBC