Home Podcasts Women in AI Research (WiAIR)
Women in AI Research (WiAIR)

Women in AI Research (WiAIR)

WiAIR 32 Episodes Aug 19, 2026

Women in AI Research (WiAIR) is a podcast dedicated to celebrating the remarkable contributions of female AI researchers from around the globe. It challenges the perception that AI research is predominantly male-driven and aims to empower early career researchers, especially women, to pursue their passion for AI. Listeners learn from women at different career stages, stay updated on the latest research and advancements, and hear powerful stories of overcoming obstacles and breaking stereotypes.

Episodes

Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss
Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss Aug 19, 2026 01:17:30 An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics th
Is Your AI Just Flattering You? Sycophancy, AI Policies, and More, with Dr. Malihe Alikhani
Is Your AI Just Flattering You? Sycophancy, AI Policies, and More, with Dr. Malihe Alikhani Jul 22, 2026 01:07:43 Only 19% of Americans say AI has actually improved their productivity - so why the gap between the hype and reality? In this episode of Women in AI Research, Dr. Malihe Alikhani (Northeastern University, Contextual AI Lab) unpacks the hidden failures in how we build and deploy AI: why sycophancy is really a collapse of alignment, why bigger models aren't better aligned, and why "thin&quot
Can Language Alone Create Intelligence? Insights from Neuroscience and AI, with Dr. Anna Ivanova
Can Language Alone Create Intelligence? Insights from Neuroscience and AI, with Dr. Anna Ivanova Jun 17, 2026 01:11:51 Do large language models truly understand language—or are they sophisticated pattern matchers?In this conversation, Dr. Anna Ivanova (Asst. Prof. at Georgia Tech) explores one of the important questions in AI: the relationship between language, thought, and intelligence. Drawing from neuroscience, cognitive science, and AI research, Anna explains why language understanding is harder to define than
100% Jailbreak Success? The Hard Truth About AI Safety, with Dr. Saadia Gabriel (Part 2)
100% Jailbreak Success? The Hard Truth About AI Safety, with Dr. Saadia Gabriel (Part 2) Apr 17, 2026 00:33:51 What actually happens when AI systems fail in the real world?In this final part of our conversation with Saadia Gabriel (UCLA), we unpack one of the most urgent challenges in modern AI: why even the most advanced models remain vulnerable to manipulation - and what that means for safety, fairness, and society.From multi-turn jailbreaking attacks with near 100% success rates to misinformation shapin
From Hate Speech to Best Paper: Building Safer AI Systems, with Dr. Saadia Gabriel (Part 1)
From Hate Speech to Best Paper: Building Safer AI Systems, with Dr. Saadia Gabriel (Part 1) Apr 15, 2026 00:29:22 What does it mean to build AI systems we can actually trust?In this first part of our conversation with Saadia Gabriel (UCLA), we explore the deeply personal and technical journey behind her work on AI safety, misuse, and responsible NLP.From experiencing targeted hate speech firsthand to receiving a best paper nomination, Saadia shares how her lived experience shaped her research — and why langua
EACL 2026: LLMs Can Hear… But Can They Reason? A New Benchmark for Audio Intelligence
EACL 2026: LLMs Can Hear… But Can They Reason? A New Benchmark for Audio Intelligence Apr 13, 2026 00:18:22 What does it actually mean for a model to understand audioPaper: https://arxiv.org/abs/2601.19673In this episode, I talk with Iwona Christop, a PhD student at Adam Mickiewicz University, about her recent EACL paper introducing ART (Audio Reasoning Tasks) — a new benchmark designed to evaluate whether multimodal LLMs can truly reason over audio, not just transcribe or classify it.Most existing benc
EACL 2026: LLMs Can Call Tools -- But Can They Understand Them?
EACL 2026: LLMs Can Call Tools -- But Can They Understand Them? Apr 12, 2026 00:22:12 LLM-based agents are everywhere, but most research focuses on just one step: getting the model to call the right tool. What happens after that?Paper: https://arxiv.org/abs/2510.15955In this talk, Kiran Kate (IBM Research) presents new findings from their EACL 2026 paper on a largely overlooked problem:👉 Can LLMs actually understand and use the outputs returned by tools?As tool-augmented systems be
EACL 2026: Reasoning Can Hurt LLM Safety?! Rethinking Accuracy in AI Systems
EACL 2026: Reasoning Can Hurt LLM Safety?! Rethinking Accuracy in AI Systems Apr 10, 2026 00:21:42 In this episode of #WiAIRpodcast, we dive into a subtle but critical question: Does adding reasoning actually make LLMs safer and more reliable?Paper: https://arxiv.org/abs/2510.21049Atoosa Chegini (University of Maryland, Apple) presents Reasoning's Razor (EACL 2026), where she and her collaborators examine how reasoning impacts high-stakes binary classification tasks, including safety filter
EACL 2026: Why LLMs Hallucinate, and How to Make Them Say "I Don't Know"
EACL 2026: Why LLMs Hallucinate, and How to Make Them Say "I Don't Know" Apr 3, 2026 00:15:00 LLMs are notoriously overconfident, but can we teach them to admit uncertainty?In this episode, Maor Juliet Lavi (Tel Aviv University) presents her EACL 2026 paper on Detecting Unanswerability in Large Language Models with Linear Directions.Paper: https://arxiv.org/abs/2509.22449We cover:Why prompt-based fixes for hallucinations aren’t enoughHow “unanswerability” emerges inside model representatio
EACL 2026: From Paraphrases to Diagnostics: A Fine-Grained Framework for LLM Auditing
EACL 2026: From Paraphrases to Diagnostics: A Fine-Grained Framework for LLM Auditing Apr 1, 2026 00:16:53 LLMs often give different answers to the same question, just phrased differently. But how do we measure and understand this behaviour rigorously?In this episode of the #WiAIRpodcast, Cléa Chataigner (Mila, McGill) presents AUGMENT, a user-grounded, controlled paraphrasing framework for auditing prompt sensitivity in large language models, accepted as an oral at EACL 2026.Paper: https://arxiv.org/a
EACL 2026: You're Using Persona Prompting Wrong
EACL 2026: You're Using Persona Prompting Wrong Mar 30, 2026 00:15:56 How much control do persona prompts actually give us over LLM behaviour?In this episode of #WiAIRpodcast, Jing Yang (TU Berlin) speaks about the study on persona prompting in socially sensitive tasks, including hate speech detection, sentiment analysis, and commonsense reasoning.Paper: https://arxiv.org/abs/2601.20757The paper takes a closer look at a common assumption: that adding demographic or
EACL 2026: Why Your LLM Needs Math to Think Better
EACL 2026: Why Your LLM Needs Math to Think Better Mar 27, 2026 00:19:22 What if adding math data actually improves reasoning in non-math tasks?In this episode of #WiAIRpodcast, Syeda from the Carnegie Mellon University and NVIDIA presents Nemotron CrossThink (EACL 2026 Oral) - a framework that challenges how we train LLMs for reasoning.Paper: https://arxiv.org/abs/2504.13941Key insights you don’t want to miss:Why multi-domain training beats specialized modelsThe surpr

Recommended