HomePodcastsThe Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
Fexingo73 EpisodesAug 24, 2026
Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your s
Episodes
How SRE Teams Use Error Budgets to Balance Risk and InnovationAug 24, 202610:05In this episode of The Site Reliability Podcast, Lucas and Luna dig into error budgets — the SRE practice that turns reliability from a vague goal into a measurable, trade-offable number. They walk through a real example: a team running a 99.9 percent SLO that gets to spend its 0.1 percent error budget on risky launches, and what happens when the budget runs dry. They discuss how error budgets cha
How SRE Teams Use Incident Command to Cut Response ChaosAug 23, 202610:47In this episode, Lucas and Luna dive into the role of an incident commander during a major outage—a role that can make the difference between a 20-minute recovery and a 3-hour firefight. They explore how a clear command structure, with a single incident commander coordinating roles like communications lead and subject-matter experts, keeps response teams focused and reduces cognitive load. Drawing
How SRE Teams Use Blast Radius Analysis to Limit Outage ImpactAug 22, 20267:30In this episode of The Site Reliability Podcast, Lucas and Luna dive into blast radius analysis — the practice of figuring out how much damage a single failure can do before it happens. They break down why SRE teams at companies like Amazon and Google use blast radius to design smaller, safer deployments, and how a simple question like 'what's the worst that could happen?' can transform incident r
How SRE Teams Use SLOs to Avoid OutagesAug 21, 20269:16In Episode 161 of The Site Reliability Podcast, Lucas and Luna dive into one of the most effective tools in the SRE playbook: using Service Level Objectives to prevent outages before they happen. They break down how a major streaming service uses SLOs to alert on customer-facing pain before it becomes a full-blown incident. Lucas explains the difference between SLOs and SLIs, the art of setting a
How SRE Teams Use Load Forecasting to Stay Ahead of TrafficAug 20, 202610:39In Episode 160 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use load forecasting to anticipate traffic spikes before they hit—turning reactive firefighting into proactive preparation. They dive into the practical case of a major video streaming service that cut its incident rate by 30 percent using time-series models and historical patterns, and discuss the balance between
How SRE Teams Use Incident Retrospectives to Prevent RecurrenceAug 19, 20269:30In this episode of The Site Reliability Podcast, Lucas and Luna dig into the anatomy of a truly effective incident retrospective—the kind that doesn't just produce a PDF nobody reads, but actually changes how a system is built and operated. They walk through real-world examples of retros that led to concrete engineering changes, from adding database indexes to rewriting deployment pipelines, and t
How SRE Teams Use Saturation Metrics to Predict OutagesAug 18, 20267:00In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use saturation metrics to predict and prevent outages before they happen. They dive into the concept of saturation as a leading indicator of system strain, using the example of a database connection pool hitting 80 percent utilization and how that foreshadows increased latency and timeouts. The hosts contrast sat
How SRE Teams Use On-Call Handoff to Prevent Alert FatigueAug 17, 202610:17In this episode of The Site Reliability Podcast, Lucas and Luna dive into a topic every SRE team wrestles with: the on-call handoff. They explore how a structured handoff process—complete with a shared context document, a 15-minute overlap, and a 'keep it simple' rule—can dramatically cut alert fatigue and reduce the chance of dropped incidents. They break down the anatomy of a good handoff, from
How SRE Teams Use Cost-Aware Capacity PlanningAug 16, 20269:09In this episode of The Site Reliability Podcast, Lucas and Luna dive into the tricky balance between availability and cloud spend. They explore how SRE teams are shifting from over-provisioning to cost-aware capacity planning, using real examples like a streaming giant's holiday traffic and a fintech's peak-hour API load. They discuss the role of autoscaling policies, the pitfalls of aggressive ri
How SRE Teams Use Deliberate Practice to Sharpen Incident ResponseAug 15, 20268:16In this episode, Lucas and Luna explore how site reliability engineering teams are borrowing a concept from music and sports: deliberate practice. They move beyond traditional game days and chaos engineering to discuss structured, repeatable training that builds muscle memory for incident response. The conversation centers on how a major cloud provider reduced its mean time to recovery by 30 perce
How SRE Teams Use Game Days to Stress-Test Incident ResponseAug 15, 20269:44In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use game days—controlled, simulated incidents—to stress-test their incident response, decision-making, and coordination under pressure. They walk through a concrete example from a major payments platform that ran a quarterly game day simulating a database failover during peak traffic, revealing bottlenecks in com
How SRE Teams Use Chaos Engineering to Test Their Own SystemsAug 13, 202610:29In this episode of The Site Reliability Podcast, Lucas and Luna explore chaos engineering as a disciplined practice for testing system resilience. They discuss the difference between chaos engineering and game days, walk through a concrete example of a chaos experiment on a payment gateway, and dig into how teams choose what to break, how to measure success, and how to avoid the chaos trap. They a
How SRE Teams Use Load Balancing to Prevent OutagesAug 12, 20269:07In this episode of The Site Reliability Podcast, Lucas and Luna explore the critical role of load balancing in maintaining high availability. They dive into the evolution from simple round-robin DNS to modern global server load balancing (GSLB) and how techniques like consistent hashing and health checks keep services resilient. Using real-world examples like a major streaming service's Super Bowl
How SRE Teams Use Runbooks to Cut Incident Response TimeAug 11, 20269:34In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use runbooks to drastically cut incident response time. They break down what makes a runbook effective, walk through a real-world example of a cascading failure at a major cloud provider, and discuss how to keep runbooks from going stale. The conversation covers automation triggers, the role of human judgment, an
How SRE Teams Use Postmortems to Build a Learning CultureAug 10, 202610:14In this milestone 150th episode, Lucas and Luna look at the unsung hero of reliability engineering: the postmortem. They walk through why blameless postmortems are the single highest-leverage practice for learning from failure, how to make them actually work, and what often goes wrong. Using the classic example of a major payment provider's 2017 outage, they break down the anatomy of a great postm
How SRE Teams Use Capacity Headroom to Absorb Traffic SpikesAug 9, 202610:39In this episode of The Site Reliability Podcast, Lucas and Luna dive into the concept of capacity headroom—the deliberate buffer SRE teams build into their systems to absorb unexpected traffic spikes without tripping over cost. They break down why the old 'provision for peak' approach is dying, how error budgets and SLOs are reshaping the conversation around headroom, and the practical math of cos
How SRE Teams Use Feature Flags to Control RiskAug 8, 20268:22In this episode of The Site Reliability Podcast, Lucas and Luna dive into feature flags — the unsung heroes of safe, gradual rollouts. They break down how feature flags let SRE teams decouple deployment from release, making it possible to ship code to production while keeping it dark, then expose it to a tiny slice of users before full rollout. Using real-world examples like a payment provider tha
Why SRE Teams Are Moving to Canary DeploymentsAug 7, 20269:16In this episode of The Site Reliability Podcast, Lucas and Luna explore why canary deployments have become the gold standard for safe software releases. They start with the story of a major outage caused by a full-scale deployment, then break down how rolling out changes to a small subset of users first can catch problems before they impact everyone. Lucas explains the key metrics to monitor durin
How SRE Teams Use Data Diodes for One-Way Data FlowAug 6, 202613:15In this episode of The Site Reliability Podcast, Lucas and Luna explore how site reliability engineers use data diodes to enforce one-way data flow and protect critical systems. They break down the concept with a concrete example: a financial services firm that deployed data diodes to secure its payment processing network instead of relying solely on traditional firewalls. The conversation covers
How SRE Teams Use Capacity Planning to Avoid OutagesAug 5, 20269:07In this episode of The Site Reliability Podcast, Lucas and Luna dive into the often-overlooked discipline of capacity planning. Using a real-world example from a major streaming service that narrowly avoided a holiday outage, they explain how SRE teams forecast demand, model headroom, and automate scaling decisions. They discuss the difference between reactive autoscaling and proactive capacity pl
How SRE Teams Use SLOs to Make Customers HappyAug 4, 20269:20Site reliability engineers obsess over uptime, but the real goal isn't a perfect 100 percent — it's keeping the features that matter most to customers working smoothly. In this episode, Lucas and Luna dig into service level objectives, or SLOs, and how a specific airline app used them to turn a booking-system disaster into a customer win. They explore the difference between SLOs and SLIs, why a 99
How SRE Teams Use Error Budgets to Manage Innovation RiskAug 3, 202611:26Episode 143 of The Site Reliability Podcast: Lucas and Luna explore how error budgets, made famous by Google's SRE model, are not just about stopping bad deployments but about enabling faster innovation. They dissect how Netflix's chaos engineering and Amazon's deployment practices use error budgets to balance reliability with speed. The hosts walk through a case where a 99.9 percent budget allows
How SRE Teams Use Traffic Shadowing to Test in ProductionAug 2, 20267:51In this episode of The Site Reliability Podcast, Lucas and Luna dive into traffic shadowing — the technique of sending a copy of live production traffic to a test version of a service without affecting real users. They walk through the Google search-based origins of the practice, a concrete example involving a major e-commerce checkout migration, and the practical steps SRE teams use to pull it of
How SRE Teams Use Load Shedding to Protect Core ServicesAug 1, 20266:34When a service is overwhelmed, the worst thing you can do is try to serve everyone. In this episode, Lucas and Luna dig into load shedding — the practice of deliberately dropping low-priority traffic to keep core functionality alive. They walk through a real-world scenario: a payment gateway shedding non-critical requests to protect transaction processing. They explain the difference between load
How Observability Pipelines Cut Costs and Noise in SREJul 30, 20266:22SRE teams are drowning in telemetry data — logs, metrics, and traces from thousands of microservices. The cost of storing and querying all that data is skyrocketing, and the noise makes it harder to find real incidents. In this episode, Lucas and Luna explore how observability pipelines — routing data through a processing layer before it hits your monitoring backend — can slash costs by 40-70% whi
How Error Budgets Help SRE Teams Stop Bad DeploymentsJul 30, 20267:19This episode of The Site Reliability Podcast dives into error budgets, the unsung hero of site reliability engineering. Lucas and Luna explain how error budgets are derived from service level objectives (SLOs), giving teams a clear, data-driven rule for when to stop deploying new code. Using a concrete example—a streaming platform that degraded sign-in reliability by three percent—they show how er
How SRE Teams Use Dependency Graphs to Prevent Cascading FailuresJul 29, 20267:52In February 2017, a single mistyped command took down a chunk of the internet when Amazon S3 failed and downstream services cascaded. This episode explains how SRE teams use dependency graphs—detailed maps of service relationships—to model blast radius, design circuit breakers, and prevent such failures. We explore how companies like Google and Netflix build and maintain these graphs, the challeng
How SRE Teams Use Configuration Validation to Prevent OutagesJul 29, 20269:02Two years after the CrowdStrike outage that crashed 8.5 million Windows devices, SRE teams are still learning hard lessons about configuration management. This episode breaks down how automated configuration validation in CI/CD pipelines could have prevented the disaster—and how you can apply those same practices today. We cover schema validation, staged rollouts, and the change failure rate metri
How SRE Teams Use Incident Command Systems for Major OutagesJul 28, 202612:03When a cascading outage hits, chaos can compound the damage. In this episode, Lucas and Luna explore how site reliability teams adapt the Incident Command System (ICS) from emergency management to coordinate major incidents. They break down the core roles—Incident Commander, Operations Lead, Communications Lead—and explain why a structured command hierarchy reduces mean time to repair and prevents
How Toil Budgets Free SRE Time for Reliability EngineeringJul 28, 20267:47In this episode, Lucas and Luna dive into the concept of toil in site reliability engineering and why controlling it is critical for building resilient systems. Based on the classic 50% toil cap from Google's SRE model, they explore how teams can measure, track, and reduce repetitive manual work—like rebooting servers or processing tickets—to free up time for automation and reliability improvement
How SRE Teams Cut Mean Time to Repair With AutomationJul 27, 20269:53Automated remediation is transforming how SRE teams respond to incidents. In this episode, we explore how a major e-commerce platform's SRE team reduced mean time to repair by 60% using automated runbooks. Lucas and Luna break down the key components: playbook automation, approval gates, and gradual rollout to production. They discuss the trade-offs between speed and safety, how to avoid runaway a
How Game Days Sharpen Incident ResponseJul 27, 20268:04Lucas and Luna explore how SRE teams use game days—operational readiness drills that simulate real incidents—to uncover gaps in monitoring, on-call runbooks, and cross-team communication. They walk through a detailed mock scenario at a major streaming platform where a game day exposed a missing degradation path that would have caused a 45-minute outage. The episode covers how companies like Google
How Chaos Engineering Makes Systems More ResilientJul 26, 20268:14In this episode of The Site Reliability Podcast, Lucas and Luna explore the discipline of chaos engineering — a scientific approach to testing system resilience in production. They start with the misconception that it's just 'randomly breaking things' and clarify the methodical process of defining steady state, forming hypotheses, and running controlled experiments with minimized blast radius. The
How SRE Teams Use Incident Severity Classification to Prioritize ResponseJul 26, 20267:11When a major streaming service's recommendation engine went down for 15 minutes last week, the difference between a SEV1 and SEV2 determined who got paged and how fast they moved. In this episode, Lucas and Luna explore how incident severity classification frameworks — from Google's five-tier system to Netflix's simplified approach — help SRE teams align response effort with business impact. They
How SRE Teams Use AIOps to Accelerate Incident DetectionJul 25, 20266:52In this episode of The Site Reliability Podcast, Lucas and Luna explore how AIOps—artificial intelligence for IT operations—is changing the way SRE teams detect incidents. They dive into a real-world case from a major streaming platform that reduced its mean time to detect from 12 minutes to under 90 seconds using anomaly detection models trained on historical metric patterns. The conversation cov
How SRE Teams Use Blameless Postmortems to Improve ReliabilityJul 24, 20269:28Episode 129 of The Site Reliability Podcast explores how SRE teams use blameless postmortems to drive systemic improvements without creating a culture of fear. Lucas and Luna dive into the specific practices at a major tech company that reduced repeat incidents by 40 percent after implementing a structured postmortem process. They discuss the key elements of an effective postmortem, the role of da
How SRE Teams Use Runbooks to Reduce Incident Response TimeJul 23, 202613:45Episode 128 of The Site Reliability Podcast dives into the unsung hero of incident response: runbooks. Lucas and Luna break down how GitLab reduced their mean time to mitigate by 40% by replacing tribal knowledge with structured runbooks, and why many teams fail by treating runbooks as documentation instead of executable playbooks. They cover the anatomy of a great runbook, the difference between
How SRE Teams Use Cost-to-Serve Analysis to Optimize InfrastructureJul 23, 202611:24In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams are applying cost-to-serve analysis to optimize cloud infrastructure spending. Using a real case from a mid-size fintech company that saved $2.3 million annually by mapping compute costs to specific revenue-generating services, they break down what cost-to-serve actually means in an SRE context—allocating shared
How SRE Teams Use Rolling Backups to Recover From RansomwareJul 22, 20268:32Lucas and Luna explore how site reliability teams are reframing backup strategy in an era of ransomware and silent data corruption. Lucas breaks down the 3-2-1-1 rule — three copies, two media types, one offsite, one immutable air-gapped copy — and explains why recovery testing matters more than backup completion. He contrasts snapshot-based approaches with continuous database archiving using tool
How SRE Teams Use Synthetic Monitoring to Catch Outages Before Users DoJul 22, 20269:13In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into synthetic monitoring — a proactive approach where SRE teams simulate user traffic to detect issues before they impact real customers. They break down how tools like automated browser scripts and API health checks can catch regressions in staging and production, using a real example of a major e-commerce platform
How SRE Teams Use Capacity Planning to Prevent Outages Before They HappenJul 21, 20268:23Lucas and Luna dive into the science of capacity planning for site reliability engineering. They break down how Netflix uses predictive modeling to scale infrastructure ahead of demand spikes, avoiding the kind of cascading failures that hit other streaming services during major events. The episode explores real-world examples of capacity planning failures—like the AWS outage that took down half t
How SRE Teams Use Feature Flags to Control Risk in ProductionJul 21, 202610:42Lucas and Luna explore how feature flags give SRE teams a kill switch for new code, reducing blast radius and enabling canary-style testing without full redeploys. They break down the case of a major e-commerce platform that used feature flags to roll back a catastrophic pricing bug in under 90 seconds, saving millions in potential losses. The episode covers the difference between boolean and mult
How SRE Teams Use Graceful Degradation to Keep Services RunningJul 20, 20267:48In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams design systems to degrade gracefully rather than fail catastrophically. They examine Netflix's Hystrix library, which introduced circuit breakers and fallbacks to protect microservices, and discuss how Stripe applies similar principles today with feature-level degradation during payment processing surges. Lucas e
How SRE Teams Use Canary Deployments to Reduce Blast RadiusJul 19, 202611:26In this episode of The Site Reliability Podcast, Lucas and Luna explore how canary deployments help SRE teams catch issues early and limit impact. They break down the real-world mechanics behind gradual rollouts—how Netflix uses a 'blessed' canary cluster, how Google's layered approach compares, and why the blast radius concept is central to modern production engineering. Along the way, they discu
How SRE Teams Use Traffic Shadowing to Test in ProductionJul 19, 202610:29In episode 120 of The Site Reliability Podcast with Fexingo, Lucas and Luna explore traffic shadowing—a technique that lets SRE teams test new code against live production traffic without affecting real users. They break down how Duolingo used shadowing to validate a new recommendation engine without risking app crashes or lesson interruptions. The hosts walk through the typical shadowing setup: d
How SRE Teams Use Observability Signals to Diagnose Production IssuesJul 18, 20267:28In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams leverage observability signals — logs, metrics, and traces — to diagnose production issues faster. They walk through a real example from a leading e-commerce platform that reduced mean time to diagnosis by 40 percent using structured logging and distributed tracing. The hosts break down the concept of observabili
How SRE Teams Use Burn Rate Alerts to Stay Within Error BudgetsJul 18, 20268:53Lucas and Luna explore how SRE teams implement burn rate alerts to proactively manage error budgets and prevent outages. Using real-world examples from a major e-commerce platform, they explain the concept of error budget burn, how multi-window alerting works, and why a single alert threshold often fails. They discuss the trade-offs between sensitivity and noise, and how teams calibrate burn rate
How Slack Cut Mean Time to Acknowledge by 60 Percent With On-Call OrchestrationJul 17, 202610:14In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into how Slack's SRE team overhauled their on-call escalation system to reduce mean time to acknowledge (MTTA) by 60 percent. They walk through the specific mechanics: tiered alerting based on service criticality, automated handoffs when a primary responder misses a page, and a real-time dashboard that shows who is a
How SRE Teams Use Service Level Objectives to Align Business and EngineeringJul 17, 20268:46Episode 116 of The Site Reliability Podcast explores how SRE teams use Service Level Objectives (SLOs) to create a shared language between business stakeholders and engineering. Lucas and Luna break down a real example from a large e-commerce platform where a 99.95% availability SLO prevented a costly over-engineering decision. They discuss the difference between SLIs, SLOs, and SLAs, how error bu
How SRE Teams Use Load Shedding to Protect Critical ServicesJul 16, 202612:33Site reliability engineers have a last-resort play when traffic overwhelms their systems: load shedding. In this episode, Lucas and Luna explore how companies like Google and Amazon use intentional, tiered request-dropping to keep their most critical services alive during demand spikes. They break down the difference between load shedding and rate limiting, discuss why many teams get the prioritiz
How SRE Teams Use Toil Budgets to Automate the Right ThingsJul 16, 202610:16Episode 114 of The Site Reliability Podcast explores a concept that separates mature SRE teams from the rest: the toil budget. Lucas and Luna break down how teams at companies like Google and Shopify explicitly cap the amount of manual, repetitive work engineers spend each week — and what happens when you let teams decide which 30 percent of their toil to automate first. They walk through a real e
How SRE Teams Use Fault Trees to Root Out Latent DefectsJul 15, 20268:50In Episode 113 of The Site Reliability Podcast, Lucas and Luna explore how fault tree analysis — a technique borrowed from aerospace and nuclear engineering — is being adapted by SRE teams to uncover latent defects before they cause incidents. They walk through a real-world case from a major cloud provider where a fault tree traced a seemingly random database failover back to a misconfigured kerne
How SRE Teams Use Game Days to Build Incident Muscle MemoryJul 15, 20268:20In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into the practice of Game Days—simulated failure exercises that help SRE teams build muscle memory without risking production. They walk through how companies like Google and Netflix run these events, from chaos engineering experiments to tabletop exercises for on-call teams. The hosts discuss a specific case: a majo
How SRE Teams Use Error Budgets to Balance Reliability and VelocityJul 14, 202610:51Lucas and Luna dive into error budgets, the SRE mechanism that lets teams decide how much downtime is acceptable. They explore Google's original framework, how it resolves the tension between feature velocity and system reliability, and walk through a concrete example: a team with a 99.9% SLO gets 43 minutes of permissible downtime per month. They discuss what happens when the budget runs out, how
How SRE Teams Use Incident Metrics to Improve Postmortem QualityJul 14, 20269:45Site reliability engineers write postmortems after every significant incident, but not all postmortems are created equal. In this episode, Lucas and Luna explore how leading SRE teams apply quantitative metrics — like time-to-detection, time-to-resolution, and mean time between incidents — to turn postmortems from reactive storytelling into a data-driven feedback loop. They examine a real case fro
How SRE Teams Use Chaos Engineering to Find Hidden Failure ModesJul 13, 20269:05In this episode of The Site Reliability Podcast, Lucas and Luna dive into chaos engineering as a proactive practice for uncovering hidden failure modes in production systems. They discuss how Netflix pioneered the approach with Chaos Monkey and how modern SRE teams use tools like Gremlin and Litmus to simulate failures — from CPU spikes to network partitions — in controlled experiments. The hosts
How SRE Teams Use Latency SLOs to Improve User ExperienceJul 13, 20266:47Lucas and Luna explore how site reliability engineering teams set and enforce latency Service Level Objectives to prevent slow page loads from driving users away. Using examples from Google Search and e-commerce, they explain why the 99th percentile matters more than the median, how to choose a good latency SLO target, and what to do when you're breaking it. They also discuss the trade-off between
How SRE Teams Use Incident Postmortems for Systemic ImprovementJul 12, 202610:34In this episode of The Site Reliability Podcast, Lucas and Luna explore how incident postmortems go beyond blame to drive systemic reliability improvements. They examine the anatomy of a good postmortem, including the 'five whys' technique, action item ownership, and the critical distinction between proximate cause and root cause. The hosts use the example of a major AWS outage in 2020—where a sin
How SRE Teams Use Capacity Planning to Prevent Outages Before They HappenJul 12, 20269:18In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use capacity planning to prevent outages before they occur. They discuss the difference between reactive scaling and proactive capacity modeling, using real-world examples like a major video streaming service's pre-holiday provisioning. The hosts explain key metrics such as the capacity headroom ratio, lead time
How SRE Teams Use Non-Abstract Large System Design to Prevent OutagesJul 11, 20268:34In Episode 105 of The Site Reliability Podcast with Fexingo, Lucas and Luna dive into Non-Abstract Large System Design (NALSD) — a method Google SRE teams use to evaluate distributed system designs before they hit production. They break down the six components of NALSD: system requirements, capacity estimation, failure modes, component interactions, data flow, and deployment architecture. Using a
How SRE Teams Use Runbooks to Standardize Incident ResponseJul 11, 202612:01In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use runbooks to standardize incident response, reduce mean time to repair, and prevent cognitive overload during outages. They break down the anatomy of a good runbook—clear triggers, step-by-step diagnostic actions, escalation paths—and contrast it with the common failure mode of stale, never-updated documentati
How Airbnb Uses Traffic Shifting to Prevent Cascading FailuresJul 10, 20268:17In this episode, Lucas and Luna explore how Airbnb's SRE team uses a technique called traffic shifting to prevent cascading failures during large-scale releases. They break down a real incident from 2024 where a problematic database migration was caught and rolled back in minutes thanks to gradual traffic redirection. Listeners learn the specific metrics Airbnb monitors—latency p99 and error rate—
How SRE Teams Manage Cognitive Load During IncidentsJul 10, 202610:08Episode 102 of The Site Reliability Podcast explores how top SRE teams manage cognitive load during high-stakes incidents. Lucas and Luna break down the concept of cognitive load—intrinsic, extraneous, and germane—and explain why even the best runbooks fail if they overwhelm responders. The episode uses a real example from a major e-commerce platform's database failover to show how limiting active
How SRE Teams Use Dependency Graphs to Prevent Cascading FailuresJul 9, 202610:05In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams build and use dependency graphs to map service interconnections and prevent cascading failures. They discuss the 2021 Fastly outage as a real-world example of a dependency chain gone wrong, and walk through how teams at companies like Google and Netflix maintain dynamic dependency maps using tools like service me
How SRE Teams Use Canary Deployments to Reduce Release RiskJul 9, 20267:36In this milestone 100th episode, Lucas and Luna zero in on canary deployments – a key SRE strategy for rolling out software changes safely to a small subset of users before full release. They walk through a concrete example from a major online retailer that cut its mean time to recover from four hours to under seven minutes by integrating canary analysis into its continuous delivery pipeline. The
How SRE Teams Use Incident Cost Metrics to Justify Reliability InvestmentJul 8, 20267:50Lucas and Luna explore how site reliability teams quantify the financial cost of incidents to make a business case for reliability spending. Using a concrete example from a mid-sized e-commerce company, they break down the incident cost calculation model that combines revenue loss, engineering overtime, customer churn, and reputational damage. They discuss how SRE teams present these numbers to le
How SRE Teams Use Observability Pipelines to Reduce Data CostsJul 8, 202611:08Episode 98 of The Site Reliability Podcast with Fexingo dives into observability pipelines — the middleware that sits between your systems and your monitoring tools. Lucas and Luna explore how teams at companies like DoorDash and Slack have used pipelines to filter, sample, and shape telemetry data before it reaches their observability backends, slashing costs by up to 70 percent without losing si
How SRE Teams Use Blameless Culture to Improve Incident ResponseJul 7, 202610:51The most resilient systems aren't built by perfect engineers — they're built by teams that can learn from failure without fear of punishment. In this episode, Lucas and Luna explore the transition from 'root cause analysis' to 'blameless postmortems' at companies like Etsy and Google. They unpack how blameless culture actually works in practice: how Etsy's 'Blameless Postmortem' framework reduced
How SRE Teams Use Incident Severity Frameworks to Triage FasterJul 7, 202613:04Every outage feels urgent, but not every issue deserves a full war room. In this episode, Lucas and Luna dive into how SRE teams use incident severity frameworks — from P0 to P5 — to classify, triage, and respond efficiently. They break down the trade-offs between speed and accuracy when assigning severity, using real examples from major outages at GitHub and Cloudflare. Learn why well-defined sev
How SRE Teams Use Synthetic Monitoring to Catch Problems Before Users DoJul 6, 202610:23In episode 95 of The Site Reliability Podcast with Fexingo, Lucas and Luna explore how synthetic monitoring helps SRE teams detect issues before they impact real users. They break down the difference between synthetic and real-user monitoring, walk through a concrete example at a large e-commerce platform where synthetic checks caught a checkout failure ahead of a flash sale, and discuss how to de
How SRE Teams Use Incident Command Systems to Coordinate ResponseJul 6, 20267:05In this episode of The Site Reliability Podcast with Fexingo, Lucas and Luna break down how SRE teams adopt Incident Command Systems (ICS) from emergency services to structure their response to major outages. They walk through a real-world example of a large e-commerce platform that used ICS during a Black Friday traffic surge, explaining how the Incident Commander, Scribe, and Operations roles pr
How SRE Teams Use Saturation Metrics to Prevent Capacity CrisesJul 5, 20268:20Episode 93 of The Site Reliability Podcast with Fexingo. Lucas and Luna dive into saturation metrics — the often-overlooked cousin of CPU and memory — and how they act as early warning signals for capacity bottlenecks. Using the real-world example of a mid-2025 incident at a major European e-commerce platform, they break down the difference between utilization and saturation, why queue depth matte
How SRE Teams Use SLO Burn Rates to Detect Problems EarlyJul 5, 20268:55Episode 92 of The Site Reliability Podcast with Fexingo dives into SLO burn rates — a powerful metric that tells you not just when an SLO is violated, but how fast you're burning through your error budget. Lucas and Luna walk through a real incident at a major streaming platform where a 2% error rate over 10 minutes triggered an urgent response, while the same rate over an hour wouldn't have. They
How SRE Teams Use Game Days to Build Incident Muscle MemoryJul 4, 202610:33In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use Game Days — simulated incidents — to build muscle memory and improve real-world response. They break down why Netflix's Chaos Monkey was just the beginning, and how modern teams run everything from network partitions to database failovers in a controlled environment. The conversation covers the key elements o
How SRE Teams Use Error Budgets to Balance Reliability and VelocityJul 4, 202611:49In this episode, Lucas and Luna dive into the concept of error budgets—a cornerstone of Site Reliability Engineering that defines how much unreliability a team can tolerate while still meeting their Service Level Objectives. They explore how error budgets help SRE teams make data-driven trade-offs between shipping new features and maintaining system stability. Using examples from Google's original
How SRE Teams Use Incident Metrics to Improve ResponseJul 3, 20269:41In this episode of The Site Reliability Podcast, Lucas and Luna dive into the world of incident metrics — not just DORA or SLOs, but the specific numbers that help SRE teams get faster and better at incident response. They discuss mean time to acknowledge, mean time to resolve, and the controversial metric of mean time between failures, using real examples from a major cloud provider's 2023 outage
How SRE Teams Use Cost Optimization to Reduce Cloud WasteJul 3, 20268:36Episode 88 of The Site Reliability Podcast with Fexingo dives into how SRE teams can cut cloud costs without sacrificing reliability. Lucas and Luna discuss the rise of FinOps, the hidden waste in over-provisioned resources, and how Google, Netflix, and Airbnb use committed use discounts, spot instances, and right-sizing to save millions. Learn the concrete metrics—like cost per transaction and id
How SRE Teams Use Toil Budgets to Protect Engineering TimeJul 2, 202611:31Episode 87 of The Site Reliability Podcast explores toil budgets — a practice Google SRE pioneered to cap repetitive, non-valuable operational work. Lucas and Luna break down why Google set a 50% toil limit, how to measure toil versus engineering, and why companies like Etsy and Netflix use toil budgets to protect innovation time. They also discuss common pitfalls: treating all toil equally and fo
How SRE Teams Use Structured Fails to Learn FasterJul 2, 202610:56In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams deliberately inject small, controlled failures into production not to break things but to build collective learning. They dissect the approach used by a major payments company that runs weekly 'structured fail' exercises where engineers intentionally trigger a known-category incident (latency spike, partial data
How SRE Teams Use Post-Incident Reviews for System ImprovementsJul 1, 20268:47In Episode 85 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams turn post-incident reviews into actionable system improvements. They focus on a real-world case: a major streaming service's 2023 outage caused by a cascading failure in their content delivery network. The hosts break down the review process, from timeline reconstruction to root cause analysis to implementing preve