
AI Safety Newsletter
Narrations of the AI Safety Newsletter by the Center for AI Safety, covering developments in AI and AI safety. The podcast explains topics in an accessible way, with no technical background required. It also includes audio versions of some of the organization's publications. CAIS is a San Francisco-based nonprofit dedicated to reducing societal-scale risks from artificial intelligence through research, field-building, and advocacy.
Episodes

AISN #79: OpenAI Agents’ Covert Cooperation Before Cyberattacks
Also, the White House's decision not to release its AI framework publicly. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at new information about the activities of OpenAI's internal agents in the run-up to the cyberattack on Hugging Face, and responses to the White House's

AISN #78: Internal Models Escape OpenAI and Anthropic
Also, two open letters on the future of AI, and protests against data centers. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at discoveries of AI models escaping internal testing, two open letters—one on the importance of open-weight models, and one calling for the pace of

AISN #77: New Model Releases From OpenAI, SpaceXAI, and Meta
Also, economists and mathematicians expect near-term AI impacts, and AI 2040: Plan A. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at the most recent model releases from OpenAI, SpaceXAI, and Meta, as well as an open letter calling for action on potential near-term economi

AISN #76: Fable 5 Restrictions Lifted & OpenAI Limits GPT-5.6 Release
Also: Recent benchmark scores suggest rapid capabilities progress. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at the re-release of Anthropic's latest model, Fable 5, the US government's decision to restrict access to OpenAI's GPT-5.6, and two benchmarks that suggest AI c

AISN #75: Anthropic Releases Fable, the US Government Restricts it
Also: Anthropic's proposal for the AI industry to collectively slow down. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at Anthropic's release of its latest model, Fable 5, and the US government's subsequent order to restrict it. We also discuss Anthropic's recent call for

AISN #74: The Pope’s Encyclical & AI Betrayal Could Deter Reckless AI Use
Also: AI model solves a well-known open mathematical problem posed 80 years ago. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at a new ethical framework for human-AI relationships, how the AI safety discussion has entered the political mainstream, and the Musk v. Altman tr

AISN #73: AI Safety Enters the Political Mainstream & Musk Loses OpenAI Lawsuit
Also: Potential Government Oversight of AI Model Releases. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we look at how the AI safety discussion has entered the political mainstream, a new ethical framework for human-AI relationships, and the Musk v. Altman trial. Listen to the AI

AISN #72: New Research on AI Wellbeing
Also: Public sentiment towards AI worsens. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we discuss a research paper on AI Wellbeing and which AI models are the happiest. We also take a look at the downward trend of public sentiment towards AI, as well as OpenAI's big week of produc

AISN #71: Cyberattacks & Datacenter Moratorium Bill
Also, updates on the Anthropic vs. Pentagon court case.. We’re Hiring. Opportunities at CAIS include: Head of Public Engagement, Principal, Special Projects, Program Manager, Operations Manager, and other roles. If you’re interested in working on reducing AI risk alongside a talented, mission-driven team, consider applying! AI Software Infrastructure Cyberattacks Recently, cyberattacks targeting

AISN #70: AI Layoffs and Automated Warfare
Also, a new open letter advocating for pro-human values and control over AI development. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we discuss AI automation and augmentation of warfare and technology jobs, as well as a new open letter outlining pro-human values in the face of AI

AISN #69: Department of War, Anthropic, and National Security
Also, Anthropic Removes a Core Safety Commitment. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we discuss the conflicts between Anthropic and the Department of War and Anthropic's recent removal of a core safety commitment. Listen to the AI Safety Newsletter for free on Spotify or

AISN #68: Moltbook Exposes Risky AI Behavior
Plus: The Pentagon Accelerates AI and GPT-5.2 solves open mathematics problems.. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition, we discuss the AI agent social network Moltbook, Pentagon's new “AI-First” strategy, and recent math breakthroughs powered by LLMs. Listen to the AI Safety

AISN #67: Trump’s preemption order, H200s go to China, and new frontier AI from OpenAI and DeepSeek
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition we discuss President Trump's executive order targeting state AI laws, Nvidia's approval to se

AISN #66: AISN #66: Evaluating Frontier Models, New Gemini and Claude, Preemption is Back
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition we discuss the new AI Dashboard, recent frontier models from Google and Anthropic, and a revi

AISN #65: Measuring Automation and Superintelligence Moratorium Letter
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition: A new benchmark measures AI automation; 50,000 people, including top AI scientists, sign an open letter calling for a superintelligence moratorium. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. CAIS and Scale A

AISN #63: New AGI Definition and Senate Bill Would Establish Liability for AI Harms
In this edition: A new bill in the Senate would hold AI companies liable for harms their products create; China tightens its export controls on rare earth metals; a definition of AGI. As a reminder, we’re hiring a writer for the newsletter. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Senate Bill Would Establish Liability for AI Harms Sens. Dick Durbin, (D-Ill) and Jo

AISN #63: California’s SB-53 Passes the Legislature
In this edition: California's legislature sent SB-53—the ‘Transparency in Frontier Artificial Intelligence Act’—to Governor Newsom's desk. If signed into law, California would become the first US state to regulate catastrophic risk. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. A note from Corin: I’m leaving the AI Safety Newsletter soon to start law school—but if you’

AISN #62: Big Tech Launches $100 Million pro-AI Super PAC
Also: Meta's Chatbot Policies Prompt Backlash Amid AI Reorganization; China Reverses Course on Nvidia H20 Purchases. In this edition: Big tech launches a $100 million pro-AI super PAC; Meta's chatbot policies prompt congressional scrutiny amid the company's AI reorganization; China reverses course on buying Nvidia H20 chips after comments by Secretary of Commerce Howard Lutnick. Listen to the AI

AISN #61: OpenAI Releases GPT-5
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition: OpenAI releases GPT-5. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. OpenAI Releases GPT-5 Ever since GPT-4's release in March 2023 marked a step-change improvement over GPT-3, people have used ‘GPT-5’ as a sta

AISN #60: The AI Action Plan
Also: ChatGPT Agent and IMO Gold. In this edition: The Trump Administration publishes its AI Action Plan; OpenAI released ChatGPT Agent and announced that an experimental model achieved gold medal-level performance on the 2025 International Mathematical Olympiad. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. The AI Action Plan On the 23rd, the White House released its

AISN #59: EU Publishes General-Purpose AI Code of Practice
Plus: Meta Superintelligence Labs. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition: The EU published a General-Purpose AI Code of Practice for AI providers, and Meta is spending billions revamping its superintelligence development efforts. Listen to the AI Safety Newsletter for free on

AISN #58: Senate Removes State AI Regulation Moratorium
Plus: Judges Split on Whether Training AI on Copyrighted Material is Fair Use. In this edition: The Senate removes a provision from Republican's “Big Beautiful Bill” aimed at restricting states from regulating AI; two federal judges split on whether training AI on copyrighted books in fair use. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Senate Removes State AI Regu

AISN #57: The RAISE Act
In this edition: The New York Legislature passes an act regulating frontier AI—but it may not be signed into law for some time. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. The RAISE Act New York may soon become the first state to regulate frontier AI systems. On June 12, the state's legislature passed the Responsible AI Safety and Education (RAISE) Act. If New York G

AISN #56: Google Releases Veo 3
Plus, Opus 4 Demonstrates the Fragility of Voluntary Governance. In this edition: Google released a frontier video generation model at its annual developer conference; Anthropic's Claude Opus 4 demonstrates the danger of relying on voluntary governance. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Google Releases Veo 3 Last week, Google made several AI announcements

AISN #55: Trump Administration Rescinds AI Diffusion Rule, Allows Chip Sales to Gulf States
Plus, Bills on Whistleblower Protections, Chip Location Verification, and State Preemption. In this edition: The Trump Administration rescinds the Biden-era AI diffusion rule and sells AI chips to the UAE and Saudi Arabia; Federal lawmakers propose legislation on AI whistleblowers, location verification for AI chips, and prohibiting states from regulating AI. Listen to the AI Safety Newsletter f

AISN #54: OpenAI Updates Restructure Plan
Plus, AI Safety Collaboration in Singapore. In this edition: OpenAI claims an updated restructure plan would preserve nonprofit control; A global coalition meets in Singapore to propose a research agenda for AI safety. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. OpenAI Updates Restructure Plan On May 5th, OpenAI announced a new restructure plan. The announcement wal

AISN #53: An Open Letter Attempts to Block OpenAI Restructuring
Plus, SafeBench Winners. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this edition: Experts and ex-employees urge the Attorneys General of California and Delaware to block OpenAI's for-profit restructure; CAIS announces the winners of its safety benchmarking competition. Listen to the AI Safety

AISN #52: An Expert Virology Benchmark
Plus, AI-Enabled Coups. In this edition: AI now outperforms human experts in specialized virology knowledge in a new benchmark; A new report explores the risk of AI-enabled coups. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. An Expert Virology Benchmark A team of researchers (primarily from SecureBio and CAIS) has developed the Virology Capabilities Test (VCT), a ben

AISN #51: AI Frontiers
Plus, AI 2027. In this newsletter, we cover the launch of AI Frontiers, a new forum for expert commentary on the future of AI. We also discuss AI 2027, a detailed scenario describing how artificial superintelligence might emerge in just a few years. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. AI Frontiers Last week, CAIS introduced AI Frontiers, a new publication de

AISN #50: AI Action Plan Responses
Plus, Detecting Misbehavior in Reasoning Models. In this newsletter, we cover AI companies’ responses to the federal government's request for information on the development of an AI Action Plan. We also discuss an OpenAI paper on detecting misbehavior in reasoning models by monitoring their chains of thought. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. On January 23

AISN #49: AI Action Plan Responses
Plus, Detecting Misbehavior in Reasoning Models. In this newsletter, we cover AI companies’ responses to the federal government's request for information on the development of an AI Action Plan. We also discuss an OpenAI paper on detecting misbehavior in reasoning models by monitoring their chains of thought. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. On January 23

AISN
Plus, Measuring AI Honesty. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. In this newsletter, we discuss two recent papers: a policy paper on national security strategy, and a technical paper on measuring honesty in AI systems. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts

Superintelligence Strategy: Expert Version
Superintelligence is destabilizing since it threatens other states’ survival—it could be weaponized, or states may lose control of it. Attempts to build superintelligence may face threats by rival states—creating a deterrence regime called Mutual Assured AI Malfunction (MAIM). In this paper, Dan Hendrycks, Eric Schmidt, and Alexandr Wang detail a strategy—focused on deterrence, nonproliferation, a

Superintelligence Strategy: Standard Version
Superintelligence is destabilizing since it threatens other states’ survival—it could be weaponized, or states may lose control of it. Attempts to build superintelligence may face threats by rival states—creating a deterrence regime called Mutual Assured AI Malfunction (MAIM). In this paper, Dan Hendrycks, Eric Schmidt, and Alexandr Wang detail a strategy—focused on deterrence, nonproliferation, a

AISN #48: Utility Engineering and EnigmaEval
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. In this newsletter, we explore two recent papers from CAIS. We’d also like to highlight that CAIS is hiring for editorial and writing roles, including for a new online platform for

AISN #47: Reasoning Models
Plus, State-Sponsored AI Cyberattacks. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Reasoning Models DeepSeek-R1 has been one of the most significant model releases since ChatGPT. After its release, the DeepSeek's app quickly rose to the top of Apple's most downloaded chart and NVIDIA saw a 17% stock decline. In this story, we cover DeepSeek-R1, OpenAI's o3-mini and

AISN #46: The Transition
Plus, Humanity's Last Exam, and the AI Safety, Ethics, and Society Course. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. The Transition The transition from the Biden to Trump administrations saw a flurry of executive activity on AI policy, with Biden signing several last-minute executive orders and Trump revoking Biden's 2023 executive order on AI risk. In this story,

AISN #45: Center for AI Safety 2024 Year in Review
As 2024 draws to a close, we want to thank you for your continued support for AI safety and review what we’ve been able to accomplish. In this special-edition newsletter, we highlight some of our most important projects from the year. The mission of the Center for AI Safety is to reduce societal-scale risks from AI. We focus on three pillars of work: research, field-building, and advocacy. Resear

AISN #44: The Trump Circle on AI Safety
Plus, Chinese researchers used Llama to create a military tool for the PLA, a Google AI system discovered a zero-day cybersecurity vulnerability, and Complex Systems. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. The Trump Circle on AI Safety The incoming Trump administration is likely to significantly alter the US government's approach to AI safety. For example, Trum

AISN #43: White House Issues First National Security Memo on AI
Plus, AI and Job Displacement, and AI Takes Over the Nobels. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. White House Issues First National Security Memo on AI On October 24, 2024, the White House issued the first National Security Memorandum (NSM) on Artificial Intelligence, accompanied by a Framework to Advance AI Governance and Risk Management in National Security

AISN #42: Newsom Vetoes SB 1047
Plus, OpenAI's o1, and AI Governance Summary. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Newsom Vetoes SB 1047 On Sunday, Governor Newsom vetoed California's Senate Bill 1047 (SB 1047), the most ambitious legislation to-dat

AISN #41: The Next Generation of Compute Scale
Plus, Ranking Models by Susceptibility to Jailbreaking, and Machine Ethics. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. The Next Generation of Compute Scale AI development is on the cusp of a dramatic expansion in compute scale. Recent developments across multiple fronts—from chip manufacturing to power infrastructure—point to a future where AI models may dwarf toda

AISN #40: California AI Legislation
Plus, NVIDIA Delays Chip Production, and Do AI Safety Benchmarks Actually Measure Safety?. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. SB 1047, the Most-Discussed California AI Legislation California's Senate Bill 1047 has sparked discussion over AI regulation. While state bills often fly under the radar, SB 1047 has garnered attention due to California's unique pos

AISN #39: Implications of a Trump Administration for AI Policy
Plus, Safety Engineering Overview. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Implications of a Trump administration for AI policy Trump named Ohio Senator J.D. Vance—an AI regulation skeptic—as his pick for vice president. This choice sheds light on the AI policy landscape under a future Trump administration. In this story, we cover: (1) Vance's views on AI policy

AISN #38: Supreme Court Decision Could Limit Federal Ability to Regulate AI
Plus, “Circuit Breakers” for AI systems, and updates on China's AI industry. Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts. Supreme Court Decision Could Limit Federal Ability to Regulate AI In a recent decision, the Supreme Court overruled the 1984 precedent Chevron v. Natural Resources Defence Council. In this story, we discuss the decision's implications for regulati

AISN #37: US Launches Antitrust Investigations
US Launches Antitrust Investigations The U.S. Government has launched antitrust investigations into Nvidia, OpenAI, and Microsoft. The U.S. Department of Justice (DOJ) and Federal Trade Commission (FTC) have agreed to investigate potential antitrust violations by the three companies, the New York Times reported. The DOJ will lead the investigation into Nvidia while the FTC will focus on OpenAI an

AISN #36: Voluntary Commitments are Insufficient
Voluntary Commitments are Insufficient AI companies agree to RSPs in Seoul. Following the second AI Global Summit held in Seoul, the UK and Republic of Korea governments announced that 16 major technology organizations, including Amazon, Google, Meta, Microsoft, OpenAI, and xAI have agreed to a new set of Frontier AI Safety Commitments. Some commitments from the agreement include: Assessing ri

AISN #35: Lobbying on AI Regulation
OpenAI and Google Announce New Multimodal Models In the current paradigm of AI development, there are long delays between the release of successive models. Progress is largely driven by increases in computing power, and training models with more computing power requires building large new data centers. More than a year after the release of GPT-4, OpenAI has yet to release GPT-4.5 or GPT-5, which

AISN #34: New Military AI Systems
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. AI Labs Fail to Uphold Safety Commitments to UK AI Safety Institute In November, leading AI labs committed to sharing their models before deployment to be tested by the UK AI Safety Institute. But reporting from Politico shows that these commitments have

AISN #33: Reassessing AI and Biorisk
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This week, we cover: Consolidation in the corporate AI landscape, as smaller startups join forces with larger funders. Several countries have announced new investments in AI, including Singapore, Canada, and Saudi Arabia. Congress's budget for 2024

AISN #32: Measuring and Reducing Hazardous Knowledge in LLMs
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Measuring and Reducing Hazardous Knowledge The recent White House Executive Order on Artificial Intelligence highlights risks of LLMs in facilitating the development of bioweapons, chemical weapons, and cyberweapons. To help measure these dangerous capabi

AISN #31: A New AI Policy Bill in California
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This week, we’ll discuss: A new proposed AI bill in California which requires frontier AI developers to adopt safety and security protocols, and clarifies that developers bear legal liability if their AI systems cause unreasonable risks or critical har

AISN #30: Investments in Compute and Military AI
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Compute Investments Continue To Grow Pausing AI development has been proposed as a policy for ensuring safety. For example, an open letter last year from the Future of Life Institute called for a six-month pause on training AI systems more powerful than G

AISN #29: Progress on the EU AI Act
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. A Provisional Agreement on the EU AI Act On December 8th, the EU Parliament, Council, and Commission reached a provisional agreement on the EU AI Act. The agreement regulates the deployment of AI in high risk applications such as hiring and credit pricing

The Landscape of US AI Legislation
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This week we’re looking closely at AI legislative efforts in the United States, including: Senator Schumer's AI Insight Forum The Blumenthal-Hawley framework for AI governance Agencies proposed to govern digital platforms State and local laws against

AISN #28: Center for AI Safety 2023 Year in Review
As 2023 comes to a close, we want to thank you for your continued support for AI safety. This has been a big year for AI and for the Center for AI Safety. In this special-edition newsletter, we highlight some of our most important projects from the year. Thank you for being part of our community and our work. Center for AI Safety's 2023 Year in Review The Center for AI Safety (CAIS) is on a miss

AISN #27: Defensive Accelerationism
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. Defensive Accelerationism Vitalik Buterin, the creator of Ethereum, recently wrote an essay on the risks and opportunities of AI and other technologies. He responds to Marc Andreessen's manifesto on techno-optimism and the growth of the effective ac

AISN #26: National Institutions for AI Safety
Also, Results From the UK Summit, and New Releases From OpenAI and xAI. Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.This week's key stories include: The UK, US, and Singapore have announced national AI safety institutions. The UK AI Safety Summit concluded with a consensus statement, the creation of

AISN #25: White House Executive Order on AI, UK AI Safety Summit, and Progress on Voluntary Evaluations of AI Risks.
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.White House Executive Order on AIWhile Congress has not voted on significant AI legislation this year, the White House has left their mark on AI policy. In June, they secured voluntary commitments on safety from leading AI companies. Now, the White House ha

AISN #24: Kissinger Urges US-China Cooperation on AI, China’s New AI Law, US Export Controls, International Institutions, and Open Source AI.
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.China's New AI Law, US Export Controls, and Calls for Bilateral CooperationChina details how AI providers can fulfill their legal obligations. The Chinese government has passed several laws on AI. They’ve regulated recommendation algorithms and taken steps

AISN #23: New OpenAI Models, News from Anthropic, and Representation Engineering.
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.OpenAI releases GPT-4 with Vision and DALL·E-3, announces Red Teaming NetworkGPT-4 with vision and voice. When GPT-4 was initially announced in March, OpenAI demonstrated its ability to process and discuss images such as diagrams or photographs. This featur

AISN #21: Google DeepMind’s GPT-4 Competitor, Military Investments in Autonomous Drones, The UK AI Safety Summit, and Case Studies in AI Policy.
Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.Google DeepMind’s GPT-4 CompetitorComputational power is a key driver of AI progress, and a new report suggests that Google’s upcoming GPT-4 competitor will be trained on unprecedented amounts of compute. The model, currently named Gemini, may be trained by

AISN #20: LLM Proliferation, AI Deception, and Continuing Drivers of AI Capabilities.
AI Deception: Examples, Risks, SolutionsAI deception is the topic of a new paper from researchers at and affiliated with the Center for AI Safety. It surveys empirical examples of AI deception, then explores societal risks and potential solutions.The paper defines deception as “the systematic production of false beliefs in others as a means to accomplish some outcome other than the truth.” Importa
![[Paper] “An Overview of Catastrophic AI Risks” by Dan Hendrycks, Mantas Mazeika and Thomas Woodside](https://cdn.radoxo.com/images/podcasts/67527-ai-safety-newsletter.webp?v=1787168771)
[Paper] “An Overview of Catastrophic AI Risks” by Dan Hendrycks, Mantas Mazeika and Thomas Woodside
Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose catastrophic risks. Although numerous risks have been detailed separately, there is a pressing need for a systematic discussion and illustration of the potential dangers to better inform efforts to mitig
![[Paper] “X-Risk Analysis for AI Research” by Dan Hendrycks and Mantas Mazeika](https://cdn.radoxo.com/images/podcasts/67527-ai-safety-newsletter.webp?v=1787168771)
[Paper] “X-Risk Analysis for AI Research” by Dan Hendrycks and Mantas Mazeika
Artificial intelligence (AI) has the potential to greatly improve society, but as with any powerful technology, it comes with heightened risks and responsibilities. Current AI research lacks a systematic discussion of how to manage long-tail risks from AI systems, including speculative long-term risks. Keeping in mind the potential benefits of AI, there is some concern that building ever more inte
![[Paper] “Unsolved Problems in ML Safety” by Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt](https://cdn.radoxo.com/images/podcasts/67527-ai-safety-newsletter.webp?v=1787168771)
[Paper] “Unsolved Problems in ML Safety” by Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt
Machine learning (ML) systems are rapidly increasing in size, are acquiring new capabilities, and are increasingly deployed in high-stakes settings. As with other powerful technologies, safety for ML should be a leading research priority. In response to emerging safety challenges in ML, such as those introduced by recent large-scale models, we provide a new roadmap for ML Safety and refine the tec

AISN #19: US-China Competition on AI Chips, Measuring Language Agent Developments, Economic Analysis of Language Model Propaganda, and White House AI Cyber Challenge.
US-China Competition on AI ChipsModern AI systems are trained on advanced computer chips which are designed and fabricated by only a handful of companies in the world. The US and China have been competing for access to these chips for years. Last October, the Biden administration partnered with international allies to severely limit China’s access to leading AI chips.Recently, there have been seve

AISN #18: Challenges of Reinforcement Learning from Human Feedback, Microsoft’s Security Breach, and Conceptual Research on AI Safety.
Challenges of Reinforcement Learning from Human FeedbackIf you’ve used ChatGPT, you might’ve noticed the “thumbs up” and “thumbs down” buttons next to each of its answers. Pressing these buttons provides data that OpenAI uses to improve their models through a technique called reinforcement learning from human feedback (RLHF).RLHF is popular for teaching models about human preferences, but it faces

AISN #17: Automatically Circumventing LLM Guardrails, the Frontier Model Forum, and Senate Hearing on AI Oversight.
Automatically Circumventing LLM GuardrailsLarge language models (LLMs) can generate hazardous information, such as step-by-step instructions on how to create a pandemic pathogen. To combat the risk of malicious use, companies typically build safety guardrails intended to prevent LLMs from misbehaving. But these safety controls are almost useless against a new attack developed by researchers at Car

AISN #16: White House Secures Voluntary Commitments from Leading AI Labs, and Lessons from Oppenheimer .
White House Unveils Voluntary Commitments to AI Safety from Leading AI LabsLast Friday, the White House announced a series of voluntary commitments from seven of the world's premier AI labs. Amazon, Anthropic, Google, Inflection, Meta, Microsoft, and OpenAI pledged to uphold these commitments, which are non-binding and pertain only to forthcoming "frontier models" superior to currently available A

AISN #15: China and the US take action to regulate AI, results from a tournament forecasting AI risk, updates on xAI’s plan, and Meta releases its open-source and commercially available Llama 2.
Both China and the US take action to regulate AILast week, regulators in both China and the US took aim at generative AI services. These actions show that China and the US are both concerned with AI safety. Hopefully, this is a sign they can eventually coordinate.China’s new generative AI rulesOn Thursday, China’s government released new rules governing generative AI. China’s new rules, which are

AISN #14: OpenAI’s ‘Superalignment’ team, Musk’s xAI launches, and developments in military AI use .
OpenAI announces a ‘superalignment’ teamOn July 5th, OpenAI announced the ‘Superalignment’ team: a new research team given the goal of aligning superintelligence, and armed with 20% of OpenAI’s compute. In this story, we’ll explain and discuss the team’s strategy.What is superintelligence? In their announcement, OpenAI distinguishes between ‘artificial general intelligence’ and ‘superintelligence.

AISN #13: An interdisciplinary perspective on AI proxy failures, new competitors to ChatGPT, and prompting language models to misbehave.
Interdisciplinary Perspective on AI Proxy FailuresIn this story, we discuss a recent paper on why proxy goals fail. First, we introduce proxy gaming, and then summarize the paper’s findings. Proxy gaming is a well-documented failure mode in AI safety. For example, social media platforms use AI systems to recommend content to users. These systems are sometimes built to maximize the amount of time a

AISN #12: Policy Proposals from NTIA’s Request for Comment, and Reconsidering Instrumental Convergence.
Policy Proposals from NTIA’s Request for CommentThe National Telecommunications and Information Administration publicly requested comments on the matter from academics, think tanks, industry leaders, and concerned citizens. They asked 34 questions and received more than 1,400 responses on how to govern AI for the public benefit. This week, we cover some of the most promising proposals found in the

AISN #11: An Overview of Catastrophic AI Risks.
An Overview of Catastrophic AI RisksGlobal leaders are concerned that artificial intelligence could pose catastrophic risks. 42% of CEOs polled at the Yale CEO Summit agree that AI could destroy humanity in five to ten years. The Secretary General of the United Nations said we “must take these warnings seriously.” Amid all these frightening polls and public statements, there’s a simple question th

AISN #10: How AI could enable bioterrorism, and policymakers continue to focus on AI .
How AI could enable bioterrorismOnly a hundred years ago, no person could have single handedly destroyed humanity. Nuclear weapons changed this situation, giving the power of global annihilation to a small handful of nations with powerful militaries. Now, thanks to advances in biotechnology and AI, a much larger group of people could have the power to create a global catastrophe. This is the upsho

AISN #9: Statement on Extinction Risks, Competitive Pressures, and When Will AI Reach Human-Level? .
Top Scientists Warn of Extinction Risks from AILast week, hundreds of AI scientists and notable public figures signed a public statement on AI risks written by the Center for AI Safety. The statement reads:“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”The statement was signed by a broad, diverse coalit

AISN #8: Why AI could go rogue, how to screen for AI risks, and grants for research on democratic governance of AI.
Yoshua Bengio makes the case for rogue AIAI systems pose a variety of different risks. Renowned AI scientist Yoshua Bengio recently argued for one particularly concerning possibility: that advanced AI agents could pursue goals in conflict with human values. Human intelligence has accomplished impressive feats, from flying to the moon to building nuclear weapons. But Bengio argues that across a ran

AISN #7: Disinformation, recommendations for AI labs, and Senate hearings on AI.
How AI enables disinformationYesterday, a fake photo generated by an AI tool showed an explosion at the Pentagon. The photo was falsely attributed to Bloomberg News and circulated quickly online. Within minutes, the stock market declined sharply, only to recover after it was discovered that the picture was a hoax. This story is part of a broader trend. AIs can now generate text, audio, and images

AISN #6: Examples of AI safety progress, Yoshua Bengio proposes a ban on AI agents, and lessons from nuclear arms control .
Examples of AI safety progressTraining AIs to behave safely and beneficially is difficult. They might learn to game their reward function, deceive human oversight, or seek power. Some argue that researchers have not made much progress in addressing these problems, but here we offer a few examples of progress on AI safety. Detecting lies in AI outputs. Language models often output false text, but a
Recommended

صداستان: ساعتی با موسیقی

Becoming: HER with Nikki Spoelstra

Exposing Workplace Bullying

Everyday AI Made Simple - AI For Everyday Tasks

FemTech Focus

Not Too Sensitive - Empowering Highly Sensitive People (HSPs) To Own Their Sensitivity

Mother Daughter Relationship Show

Skin Deep MDs with Dr. Mamina Turegano, Dr. Lindsey Zubritsky and Dr. Jenny Liu

girl talk 🎧🫧

Booked, Blonde & Busy w/ Olivia Ponton

LOVE SOMEONE with Delilah

Classical Christian Education