Home Podcasts AI Safety - Paper Digest
AI Safety - Paper Digest

AI Safety - Paper Digest

Arian Abbasi, Alan Aqrawi 13 Episodes Dec 10, 2025

This podcast breaks down the latest research and developments in AI safety, episode by episode. Each installment takes a deep dive into new cutting-edge papers, making complex ideas accessible to both experts and those who are just AI-curious. The show aims to keep listeners informed and ahead of the curve on AI security topics. Note that the podcast and its content are AI-generated, so listeners are encouraged to verify information independently.

Episodes

AI Agents: Adoption and Usage | Perplexity Comet
AI Agents: Adoption and Usage | Perplexity Comet Dec 10, 2025 00:05:02 Dive into the key findings of the first large-scale field study on the adoption, usage intensity, and use cases of general-purpose AI agents, drawing on hundreds of millions of anonymized user interactions with Perplexity’s Comet Assistant. This paper tracks the frontier shift from conversational LLM chatbots to action-oriented AI agents, which are defined as AI assistants capable of autonomously
WEF & Accenture | Advancing Responsible AI Innovation: A Playbook
WEF & Accenture | Advancing Responsible AI Innovation: A Playbook Sep 25, 2025 00:02:19 This episode of the AI Safety Paper Digest is about the World Economic Forum's new playbook on advancing responsible AI innovation. In cooperation with Accenture, the report provides a practical roadmap for turning responsible AI from an aspiration into a competitive advantage while building public trust.Link to the Report: https://www.weforum.org/publications/advancing-responsible-ai-innovati
Okay Waymo, Crash My Car! 🗣️ Testing Autonomous Vehicle Safety with Adversarial Driving Scenarios | LD-Scene
Okay Waymo, Crash My Car! 🗣️ Testing Autonomous Vehicle Safety with Adversarial Driving Scenarios | LD-Scene Aug 20, 2025 00:18:15 How can we make autonomous driving systems safer through generative AI? In this episode, we explore LD-Scene, a novel framework that combines Large Language Models (LLMs) with Latent Diffusion Models (LDMs) to create controllable, safety-critical driving scenarios. These adversarial scenarios are essential for evaluating and stress-testing autonomous vehicles, yet they’re extremely rare in real-wo
The Full LLM Glossary and Foundations
The Full LLM Glossary and Foundations Aug 4, 2025 01:28:18 Ever wanted a clear, comprehensive explanation of all the key terms related to Large Language Models (LLMs)? This episode has you covered.In this >1-hour deep-dive, we'll guide you through the essential glossary of LLM-related terms and foundational concepts, perfect for listening while driving, working, or on the go. Whether you're new to LLMs or looking to reinforce your understanding
Anthropic's Best-of-N: Cracking Frontier AI Across Modalities
Anthropic's Best-of-N: Cracking Frontier AI Across Modalities Dec 25, 2024 00:12:37 In this special christmas episode, we delve into "Best-of-N Jailbreaking," a powerful new black-box algorithm that demonstrates the vulnerabilities of cutting-edge AI systems. This approach works by sampling numerous augmented prompts - like shuffled or capitalized text - until a harmful response is elicited. Discover how Best-of-N (BoN) Jailbreaking achieves: 89% Attack Success Rates (ASR) on G
Auto-Rewards & Multi-Step RL for Diverse AI Attacks by OpenAI
Auto-Rewards & Multi-Step RL for Diverse AI Attacks by OpenAI Nov 30, 2024 00:11:17 In this episode, we explore the latest advancements in automated red teaming from OpenAI, presented in the paper "Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning." Automated red teaming has become essential for discovering rare failures and generating challenging test cases for large language models (LLMs). This paper tackles a core challenge: ho
Battle of the Scanners: Top Red Teaming Frameworks for LLMs
Battle of the Scanners: Top Red Teaming Frameworks for LLMs Nov 4, 2024 00:14:47 In this episode, we explore the findings from "Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis." As large language models (LLMs) are integrated into more applications, so do the security risks they pose, including information leaks and jailbreak attacks. This study examines four major open-source vulnerability scanners - Garak, Giskard, PyRIT, and CyberS
Watermarking LLM Output: SynthID by DeepMind
Watermarking LLM Output: SynthID by DeepMind Oct 24, 2024 00:12:57 In this episode, we delve into the groundbreaking watermarking technology presented in the paper "Scalable Watermarking for Identifying Large Language Model Outputs," published in Nature. SynthID-Text, a new watermarking scheme developed for large-scale production systems, preserves text quality while enabling high detection accuracy for synthetic content. We explore how this technology tackles th
Open Source Red Teaming: PyRIT by Microsoft
Open Source Red Teaming: PyRIT by Microsoft Oct 8, 2024 00:10:53 In this episode, we dive into PyRIT, the open-source toolkit developed by Microsoft for red teaming and security risk identification in generative AI systems. PyRIT offers a model-agnostic framework that enables red teamers to detect novel risks, harms, and jailbreaks in both single- and multi-modal AI models. We’ll explore how this cutting-edge tool is shaping the future of AI security and its pr
Jailbreaking GPT o1: STCA Attack
Jailbreaking GPT o1: STCA Attack Oct 7, 2024 00:08:32 This podcast, "Jailbreaking GPT o1, " explores how the GPT o1 series, known for its advanced "slow-thinking" abilities, can be manipulated into generating disallowed content like hate speech through a novel attack method, the Single-Turn Crescendo Attack (STCA), which effectively bypasses GPT o1's safety protocols by leveraging the AI's learned language patterns and its step-by-step reasoning proc
The Attack Atlas by IBM Research
The Attack Atlas by IBM Research Oct 5, 2024 00:11:14 This episode explores the intricate world of red-teaming generative AI models as discussed in the paper "Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI." We'll dive into the emerging vulnerabilities as LLMs are increasingly integrated into real-world applications and the evolving tactics of adversarial attacks. Our conversation will center around the "At
The Single-Turn Crescendo Attack
The Single-Turn Crescendo Attack Oct 4, 2024 00:06:45 In this episode, we examine the cutting-edge adversarial strategy presented in "Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)." Building on the multi-turn crescendo attack method, STCA escalates context within a single, expertly crafted prompt, effectively breaching the safeguards of large language models (LLMs) like never before. We discuss how this method can bypass moder

Recommended