
Women in AI Research (WiAIR)
Women in AI Research (WiAIR) is a podcast dedicated to celebrating the remarkable contributions of female AI researchers from around the globe. It challenges the perception that AI research is predominantly male-driven and aims to empower early career researchers, especially women, to pursue their passion for AI. Listeners learn from women at different career stages, stay updated on the latest research and advancements, and hear powerful stories of overcoming obstacles and breaking stereotypes.
Episodes

Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss
An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics th

Is Your AI Just Flattering You? Sycophancy, AI Policies, and More, with Dr. Malihe Alikhani
Only 19% of Americans say AI has actually improved their productivity - so why the gap between the hype and reality? In this episode of Women in AI Research, Dr. Malihe Alikhani (Northeastern University, Contextual AI Lab) unpacks the hidden failures in how we build and deploy AI: why sycophancy is really a collapse of alignment, why bigger models aren't better aligned, and why "thin"

Can Language Alone Create Intelligence? Insights from Neuroscience and AI, with Dr. Anna Ivanova
Do large language models truly understand language—or are they sophisticated pattern matchers?In this conversation, Dr. Anna Ivanova (Asst. Prof. at Georgia Tech) explores one of the important questions in AI: the relationship between language, thought, and intelligence. Drawing from neuroscience, cognitive science, and AI research, Anna explains why language understanding is harder to define than

100% Jailbreak Success? The Hard Truth About AI Safety, with Dr. Saadia Gabriel (Part 2)
What actually happens when AI systems fail in the real world?In this final part of our conversation with Saadia Gabriel (UCLA), we unpack one of the most urgent challenges in modern AI: why even the most advanced models remain vulnerable to manipulation - and what that means for safety, fairness, and society.From multi-turn jailbreaking attacks with near 100% success rates to misinformation shapin

From Hate Speech to Best Paper: Building Safer AI Systems, with Dr. Saadia Gabriel (Part 1)
What does it mean to build AI systems we can actually trust?In this first part of our conversation with Saadia Gabriel (UCLA), we explore the deeply personal and technical journey behind her work on AI safety, misuse, and responsible NLP.From experiencing targeted hate speech firsthand to receiving a best paper nomination, Saadia shares how her lived experience shaped her research — and why langua

EACL 2026: LLMs Can Hear… But Can They Reason? A New Benchmark for Audio Intelligence
What does it actually mean for a model to understand audioPaper: https://arxiv.org/abs/2601.19673In this episode, I talk with Iwona Christop, a PhD student at Adam Mickiewicz University, about her recent EACL paper introducing ART (Audio Reasoning Tasks) — a new benchmark designed to evaluate whether multimodal LLMs can truly reason over audio, not just transcribe or classify it.Most existing benc

EACL 2026: LLMs Can Call Tools -- But Can They Understand Them?
LLM-based agents are everywhere, but most research focuses on just one step: getting the model to call the right tool. What happens after that?Paper: https://arxiv.org/abs/2510.15955In this talk, Kiran Kate (IBM Research) presents new findings from their EACL 2026 paper on a largely overlooked problem:👉 Can LLMs actually understand and use the outputs returned by tools?As tool-augmented systems be

EACL 2026: Reasoning Can Hurt LLM Safety?! Rethinking Accuracy in AI Systems
In this episode of #WiAIRpodcast, we dive into a subtle but critical question: Does adding reasoning actually make LLMs safer and more reliable?Paper: https://arxiv.org/abs/2510.21049Atoosa Chegini (University of Maryland, Apple) presents Reasoning's Razor (EACL 2026), where she and her collaborators examine how reasoning impacts high-stakes binary classification tasks, including safety filter

EACL 2026: Why LLMs Hallucinate, and How to Make Them Say "I Don't Know"
LLMs are notoriously overconfident, but can we teach them to admit uncertainty?In this episode, Maor Juliet Lavi (Tel Aviv University) presents her EACL 2026 paper on Detecting Unanswerability in Large Language Models with Linear Directions.Paper: https://arxiv.org/abs/2509.22449We cover:Why prompt-based fixes for hallucinations aren’t enoughHow “unanswerability” emerges inside model representatio

EACL 2026: From Paraphrases to Diagnostics: A Fine-Grained Framework for LLM Auditing
LLMs often give different answers to the same question, just phrased differently. But how do we measure and understand this behaviour rigorously?In this episode of the #WiAIRpodcast, Cléa Chataigner (Mila, McGill) presents AUGMENT, a user-grounded, controlled paraphrasing framework for auditing prompt sensitivity in large language models, accepted as an oral at EACL 2026.Paper: https://arxiv.org/a

EACL 2026: You're Using Persona Prompting Wrong
How much control do persona prompts actually give us over LLM behaviour?In this episode of #WiAIRpodcast, Jing Yang (TU Berlin) speaks about the study on persona prompting in socially sensitive tasks, including hate speech detection, sentiment analysis, and commonsense reasoning.Paper: https://arxiv.org/abs/2601.20757The paper takes a closer look at a common assumption: that adding demographic or

EACL 2026: Why Your LLM Needs Math to Think Better
What if adding math data actually improves reasoning in non-math tasks?In this episode of #WiAIRpodcast, Syeda from the Carnegie Mellon University and NVIDIA presents Nemotron CrossThink (EACL 2026 Oral) - a framework that challenges how we train LLMs for reasoning.Paper: https://arxiv.org/abs/2504.13941Key insights you don’t want to miss:Why multi-domain training beats specialized modelsThe surpr

EACL 2026: Inference-Time Steering Is Riskier Than You Think
Are the techniques we use to control language models quietly making them less safe?Paper: https://arxiv.org/abs/2602.06256In this EACL 2026 paper, Navita Goyal (University of Maryland) challenges a core assumption behind inference-time interventions: that we can precisely steer model behaviour without unintended side effects.▶️ Key insight:While steering methods improve target behaviours (like red

EACL 2026: Do LLMs Spread Misinformation Better Than Humans?
What happens when humans and large language models try to persuade each other?In this episode of Women in AI Research, we sit down with Angana Borah (University of Michigan) to unpack her EACL 2026 paper on misinformation, persuasion, and demographic effects in human–LLM interactions.Paper: https://arxiv.org/abs/2503.02038Key insights from the paper:LLM-generated persuasion can reduce human accura

Does Liking Yellow Make You a School Bus Driver? Hidden Failures in LLMs, with Dr. Hila Gonen
In this conversation, Dr. Hila Gonen (Assistant Professor at the University of British Columbia) joins us to explore the deep insights into how large language models (LLMs) leak semantic information, behave across languages, and how researchers can uncover their root causes. Dr. Gonen shares her journey in interpreting AI systems, addressing biases, and controlling model outputs for safer, fairer

Faithfulness and Hallucinations in Reasoning Models, with Dr. Letitia Parcalabescu
Are reasoning models actually reasoning — or just producing convincing stories?Our guest in this episode of #WiAIRpodcast is Letitia Parcalabescu, the creator of the @AICoffeeBreak youtube channel. Letitia joins Jekaterina Novikova for a deep dive into the topics of faithfulness, self-consistency, hallucinations, and the reliability illusion in LLMs and multimodal reasoning models.We discuss why

AI Safety Beyond Benchmarks -- Dr. Swabha Swayamdipta on Evaluation, Personalization, and Control
As language models become more capable, the hardest questions are no longer just about performance, but about safety, interpretation, and control.In this episode of Women in AI Research, we speak with Swabha Swayamdipta, Assistant Professor of Computer Science at the University of Southern California and co-Associate Director of the USC Center for AI and Society. Swabha’s research examines how the

Do LLMs Understand Meaning? Neuroscience, Evaluation, and the Future of AI, with Dr. Maria Ryskina
Do large language models actually understand meaning — or are we over-interpreting impressive behavior?In this episode, we speak with Dr. Maria Ryskina, CIFAR AI Safety Postdoctoral Fellow at the Vector Institute for AI, whose research bridges neuroscience, cognitive science, and artificial intelligence. Together, we unpack what the brain can (and cannot) teach us about modern AI systems — and why

How Does AI Reflect Society, with Dr. Maria Antoniak
AI doesn’t just process text — it takes in our cultures, reflects our hierarchies, and can make existing power structures even stronger. In this episode of Women in AI Research, Jekaterina Novikova and Malikeh Ehgaghi speak with Dr. Maria Antoniak (Assistant Professor at the University of Colorado Boulder) about inclusivity in AI, the dynamics of cultural representation, what trust in AI really me

Multilingual AI, with Dr. Annie En-Shiun Lee
Is English just one of the languages you speak? If so, the AI tools you use might miss things that makes your voice multilingual.In this episode of Women in AI Research, Jekaterina Novikova speaks with Dr. Annie En-Shiun Lee about her work on multilingual and multicultural AI — from the widening language gap and the lack of benchmarks for underrepresented languages, to why domain-specific data mat

Why AI Doesn’t Understand Your Culture? Dr. Vered Shwartz on Cultural Bias in LLMs
Are today’s AI systems truly global — or just Western by design? 🌍In this episode of Women in AI Research, Jekaterina Novikova and Malikeh Ehgaghi speak with Dr. Vered Shwartz (Assistant Professor at UBC and CIFAR AI Chair at the Vector Institute) about the cultural blind spots in today’s large language and vision-language models.REFERENCES:Vered Shwartz Google Scholar profileBook "Lost in Au

Can We Trust AI Explanations? Dr. Ana Marasović on AI Trustworthiness, Explainability & Faithfulness
In this conversation, Jekaterina Novikova and Malikeh Ehgaghi interview Ana Marasović, an expert in AI trustworthiness, focusing on the complexities of explainability, the realities of academic research, and the dynamics of human-AI collaboration. We discuss the importance of intrinsic and extrinsic trust in AI systems, the challenges of evaluating AI performance, and the implications of synthetic

Open Science and LLMs, with Dr. Valentina Pyatkin
Can open-source large language models really outperform closed ones like Claude 3.5? 🤔In this episode of the Women in AI Research podcast, Jekaterina Novikova and Malikeh Ehghaghi engage with Valentina Pyatkin, a postdoctoral researcher at the Allen Institute for AI. We dive deep into the future of open science, LLM research, and extending model capabilities.🔑 Topics we cover:Why open-source LLMs

Unlocking LLM Reasoning, with Simeng Sophia Han
How can we go beyond accuracy to truly understand large language models?In this episode of the Women in AI Research podcast, hosts Jekaterina Novikova and Malikeh Ehghaghi sit down with Simeng Sophia Han (PhD candidate at #yaleuniversity , Research Scientist Intern at #metaai , ex #google #deepmind , ex #aws ) to explore the future of 𝐋𝐋𝐌 𝐫𝐞𝐚𝐬𝐨𝐧𝐢𝐧𝐠, 𝐞𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧, 𝐚𝐧𝐝 𝐞𝐱𝐩𝐥𝐚𝐢𝐧𝐚𝐛𝐥𝐞 𝐀𝐈.🌟 What you’ll lea

LLM Hallucinations and Machine Unlearning, with Dr. Abhilasha Ravichander
In this episode of the Women in AI Research Podcast, hosts Jekaterina Novikova and Malikeh Ehghaghi engage with Abhilasha Ravichander to discuss the complexities of LLM hallucinations, the development of factuality benchmarks, and the importance of data transparency and machine unlearning in AI. The conversation also delves into personal experiences in academia and the future directions of researc

Generalization in AI, with Dr. Dieuwke Hupkes
A must-listen episode with Dr. Dieuwke Hupkes, a research scientist at #Meta AI Research, where we dive into AI generalization, LLM robustness, and model evaluation in large language models.We explore how LLMs handle grammar and hierarchy, how they generalize across tasks and languages, and what consistency tells us about AI alignment.We also talk about Dieuwke’s journey from physics to NLP, the c

Decentralized AI, with Wanru Zhao
🔍 𝐂𝐚𝐧 𝐭𝐡𝐞 𝐟𝐮𝐭𝐮𝐫𝐞 𝐨𝐟 𝐀𝐈 𝐛𝐞 𝐝𝐞𝐜𝐞𝐧𝐭𝐫𝐚𝐥𝐢𝐳𝐞𝐝? 𝐇𝐨𝐰 𝐝𝐨 𝐰𝐞 𝐬𝐜𝐚𝐥𝐞 𝐛𝐞𝐲𝐨𝐧𝐝 𝐬𝐜𝐚𝐥𝐢𝐧𝐠 𝐥𝐚𝐰𝐬? 𝐀𝐧𝐝 𝐰𝐡𝐚𝐭 𝐝𝐨𝐞𝐬 𝐢𝐭 𝐫𝐞𝐚𝐥𝐥𝐲 𝐭𝐚𝐤𝐞 𝐭𝐨 𝐛𝐮𝐢𝐥𝐝 𝐢𝐧𝐜𝐥𝐮𝐬𝐢𝐯𝐞, 𝐦𝐮𝐥𝐭𝐢𝐥𝐢𝐧𝐠𝐮𝐚𝐥 𝐋𝐋𝐌𝐬?In this episode of the #WiAIRpodcast, Wanru Zhao discusses decentralized and collaborative AI methods, the limitations of scaling laws, fine-tuning strategies, data attribution challenges in LLMs, and multilingual learning in federated settings—all while

Interpretable AI, with Dr. Faiza Khan Khattak
How can we build AI systems that are fair, explainable, and truly responsible? In this episode of the #WiAIR podcast, we sit down with Dr. Faiza Khan Khattak, the CTO of an innovative AI startup, with a rich background in both academia and industry. From fairness in machine learning to the realities of ML deployment in healthcare, this conversation is packed with insights, real-world challenges, a

Robots with Empathy, with Dr. Angelica Lim
Dr. Angelica Lim, Assistant Professor at Simon Fraser University and Director of the SFU Rosie Lab. Can robots have feelings? In this episode, we explore the intersection of robotics, machine learning, and developmental psychology, and consider both the technical challenges and philosophical questions surrounding emotional AI. This conversation offers a glimpse into the future of human-robot inter

Responsible AI for Health, with Aparna Balagopalan
Aparna Balagopalan is a PhD student in the Department of Electrical Engineering and Computer Science (EECS) at the Massachusetts Institute of Technology.In this episode, we present the intersection of AI and healthcare. Aparna shares her research on developing fair, interpretable, and robust models for healthcare applications. We explore the unique challenges of applying AI in medical contexts, in

Bias in AI, with Amanda Cercas Curry
Dr. Amanda Cercas Curry is a researcher at CENTAI Institute, where she is working on applied NLP, fairness and evaluation.In this episode, we explore the critical issue of bias in AI systems. Amanda shares her expertise on identifying, measuring, and mitigating various forms of bias in language models and other AI applications. We discuss the social and ethical implications of biased AI, and how

Limits of Transformers, with Dr. Nouha Dziri
Nouha Dziri is an AI research scientist at the Allen Institute for AI, ex-Google DeepMind, ex-Microsoft Research.In this episode, we dive deep into the limitations of transformer models with Nouha Dziri, a research scientist at Allen Institute for AI. Nouha shares insights from her research on understanding the capabilities and constraints of LLMs. We discuss the challenges in reasoning, factualit

The new WiAIR podcast - Trailer
We are starting a new podcast!It's Women in AI Research, or simply WiAIR. Get ready for inspiring stories of leading women in AI research and their groundbreaking work. Learn from leading women in AI, hear powerful stories and join the community of AI researchers that value diversity.
Recommended

Conspiracy Files with Paige Carter

Learn English B1 with Daily News | English Listening Practice

Bad Friends

The Swerve Podcast: Obscure Topics | Conspiracy Theories

The Bread and Banter Podcast

The Church of What's Happening Now: The New Testament

Deadline: White House

English Vocabulary Help

این نقطه

Solved Murders - True Crime Stories

紐約鳥|New York Aperture

Doctor Zhivago Slow Read