
AI Safety Breakthrough
AI Safety Breakthrough, hosted by AI SafeGuard, explores the latest developments in AI safety research. Each episode examines potential risks and breakthroughs, aiming to keep AI beneficial for society. The show empowers listeners to become informed participants in conversations about AI's role. The host brings expertise from Carnegie Mellon and years in cybersecurity and AI safety.
Episodes

Navigating the New AI Security
Welcome to Agentic AI Unlocked, your deep dive into the transformative world of Agentic AI—systems combining large language models with advanced reasoning and autonomous action. These intelligent agents promise to disrupt industries, yet introduce a fundamentally new threat surface. Risks like memory poisoning, tool misuse, prompt injection, and insider threats highlight the urgent need for robust

DeepSeek: A Disruptive Force in AI
This episode explores DeepSeek, a Chinese AI startup challenging the AI landscape with its free alternative to ChatGPT. We'll examine DeepSeek's innovative architecture, including Mixture-of-Experts (MoE) and Multi-head Latent Attention (MLA), which optimize efficiency. The discussion will highlight DeepSeek's use of reinforcement learning (RL) and its impact on reasoning capabilities, as well as

VLSBench: A Visual Leakless Multimodal Safety Benchmark
Are current AI safety benchmarks for multimodal models flawed? This podcast explores the groundbreaking research behind VLSBench, a new benchmark designed to address a critical flaw in existing safety evaluations: visual safety information leakage (VSIL)We delve into how sensitive information in images is often unintentionally revealed in the accompanying text prompts, allowing models to identify

Adaptive Stress Testing for Language Model Toxicity
This episode explores ASTPrompter, a novel approach to automated red-teaming for large language models (LLMs). Unlike traditional methods that focus on simply triggering toxic outputs, ASTPrompter is designed to discover likely toxic prompts – those that could naturally emerge during regular language model use. The approach uses Adaptive Stress Testing (AST), a technique that identifies likely fai

Global Responsible AI Maturity: A Survey of 1000 Organizations
This episode dives into the critical topic of Responsible AI (RAI), exploring how organizations worldwide are grappling with the ethical and practical challenges of AI adoption. We'll be drawing insights from a comprehensive survey of 1000 organizations across 20 industries and 19 geographical regions

Ivy-VL: A Lightweight Multimodal Model for Everyday Devices
In this episode, we dive into Ivy-VL, a groundbreaking lightweight multimodal AI model released by AI Safeguard in collaboration with Carnegie Mellon University (CMU) and Stanford University. With only 3 billion parameters, Ivy-VL processes both image and text inputs to generate text outputs, offering an optimal balance of performance, speed, and efficiency. Its compact design supports deployment

Agent Bench: Evaluating LLMs as Agents
Large Language Models (LLMs) are rapidly evolving, but how do we assess their ability to act as agents in complex, real-world scenarios? Join Jenny as we explore Agent Bench, a new benchmark designed to evaluate LLMs in diverse environments, from operating systems to digital card games. We'll delve into the key findings, including the strengths and weaknesses of different LLMs and the challenges o

Hacking AI for Good: Open AI’s Red Teaming Approach
In this podcast, we delve into OpenAI's innovative approach to enhancing AI safety through red teaming—a structured process that uses both human expertise and automated systems to identify potential risks in AI models. We explore how OpenAI collaborates with external experts to test frontier models and employs automated methods to scale the discovery of model vulnerabilities. Join Jenny as we disc

Surgical Precision: PKE’s Role in AI Safety
Explore how Precision Knowledge Editing (PKE) refines AI for safety and ethical behavior in Surgical Precision: PKE’s Role in AI Safety. Join experts as we uncover the science, challenges, and breakthroughs shaping trustworthy AI. Perfect for tech enthusiasts and professionals alike, this podcast reveals how PKE ensures AI serves humanity responsibly.
Recommended

Bible Tea

TED Talks Daily

Pod Save America

Dateline NBC

صداستان: ساعتی با موسیقی

Becoming: HER with Nikki Spoelstra

Exposing Workplace Bullying

Everyday AI Made Simple - AI For Everyday Tasks

FemTech Focus

Not Too Sensitive - Empowering Highly Sensitive People (HSPs) To Own Their Sensitivity

Mother Daughter Relationship Show

Skin Deep MDs with Dr. Mamina Turegano, Dr. Lindsey Zubritsky and Dr. Jenny Liu