Home Podcasts The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering
The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

The Site Reliability Podcast with Fexingo: SRE, Uptime, and Production Engineering

Fexingo 73 Episodes Aug 24, 2026

Lucas and Luna cut through the noise around site reliability engineering to examine how real-world SRE teams balance uptime, incident response, and production change. Each episode takes a single concept — error budgets, toil automation, postmortem culture, capacity planning — and grounds it in a specific case: how a major streaming service reduced paging noise, how a payments platform rebuilt its incident command structure, or how a cloud provider manages multi-region failover. Lucas brings the numbers — latency percentiles, MTTR trends, SLO burn rates — while Luna pushes on the human and organizational trade-offs: What does a junior SRE need to know about on-call? How do you measure reliability without crushing innovation? Why do some blameless postmortems actually work? Together they treat SRE not as a certification topic but as a living practice, citing real outages, open-source tools, and engineering blogs. This show is for engineers, ops leads, and platform teams who already know the basics and want to debate the hard edges: Is 99.999% uptime always worth the cost? When should you deliberately degrade service to improve reliability? How do you design for resilience when your s

Episodes

How SRE Teams Use Error Budgets to Balance Risk and Innovation
How SRE Teams Use Error Budgets to Balance Risk and Innovation Aug 24, 2026 10:05 In this episode of The Site Reliability Podcast, Lucas and Luna dig into error budgets — the SRE practice that turns reliability from a vague goal into a measurable, trade-offable number. They walk through a real example: a team running a 99.9 percent SLO that gets to spend its 0.1 percent error budget on risky launches, and what happens when the budget runs dry. They discuss how error budgets cha
How SRE Teams Use Incident Command to Cut Response Chaos
How SRE Teams Use Incident Command to Cut Response Chaos Aug 23, 2026 10:47 In this episode, Lucas and Luna dive into the role of an incident commander during a major outage—a role that can make the difference between a 20-minute recovery and a 3-hour firefight. They explore how a clear command structure, with a single incident commander coordinating roles like communications lead and subject-matter experts, keeps response teams focused and reduces cognitive load. Drawing
How SRE Teams Use Blast Radius Analysis to Limit Outage Impact
How SRE Teams Use Blast Radius Analysis to Limit Outage Impact Aug 22, 2026 7:30 In this episode of The Site Reliability Podcast, Lucas and Luna dive into blast radius analysis — the practice of figuring out how much damage a single failure can do before it happens. They break down why SRE teams at companies like Amazon and Google use blast radius to design smaller, safer deployments, and how a simple question like 'what's the worst that could happen?' can transform incident r
How SRE Teams Use SLOs to Avoid Outages
How SRE Teams Use SLOs to Avoid Outages Aug 21, 2026 9:16 In Episode 161 of The Site Reliability Podcast, Lucas and Luna dive into one of the most effective tools in the SRE playbook: using Service Level Objectives to prevent outages before they happen. They break down how a major streaming service uses SLOs to alert on customer-facing pain before it becomes a full-blown incident. Lucas explains the difference between SLOs and SLIs, the art of setting a
How SRE Teams Use Load Forecasting to Stay Ahead of Traffic
How SRE Teams Use Load Forecasting to Stay Ahead of Traffic Aug 20, 2026 10:39 In Episode 160 of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use load forecasting to anticipate traffic spikes before they hit—turning reactive firefighting into proactive preparation. They dive into the practical case of a major video streaming service that cut its incident rate by 30 percent using time-series models and historical patterns, and discuss the balance between
How SRE Teams Use Incident Retrospectives to Prevent Recurrence
How SRE Teams Use Incident Retrospectives to Prevent Recurrence Aug 19, 2026 9:30 In this episode of The Site Reliability Podcast, Lucas and Luna dig into the anatomy of a truly effective incident retrospective—the kind that doesn't just produce a PDF nobody reads, but actually changes how a system is built and operated. They walk through real-world examples of retros that led to concrete engineering changes, from adding database indexes to rewriting deployment pipelines, and t
How SRE Teams Use Saturation Metrics to Predict Outages
How SRE Teams Use Saturation Metrics to Predict Outages Aug 18, 2026 7:00 In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use saturation metrics to predict and prevent outages before they happen. They dive into the concept of saturation as a leading indicator of system strain, using the example of a database connection pool hitting 80 percent utilization and how that foreshadows increased latency and timeouts. The hosts contrast sat
How SRE Teams Use On-Call Handoff to Prevent Alert Fatigue
How SRE Teams Use On-Call Handoff to Prevent Alert Fatigue Aug 17, 2026 10:17 In this episode of The Site Reliability Podcast, Lucas and Luna dive into a topic every SRE team wrestles with: the on-call handoff. They explore how a structured handoff process—complete with a shared context document, a 15-minute overlap, and a 'keep it simple' rule—can dramatically cut alert fatigue and reduce the chance of dropped incidents. They break down the anatomy of a good handoff, from
How SRE Teams Use Cost-Aware Capacity Planning
How SRE Teams Use Cost-Aware Capacity Planning Aug 16, 2026 9:09 In this episode of The Site Reliability Podcast, Lucas and Luna dive into the tricky balance between availability and cloud spend. They explore how SRE teams are shifting from over-provisioning to cost-aware capacity planning, using real examples like a streaming giant's holiday traffic and a fintech's peak-hour API load. They discuss the role of autoscaling policies, the pitfalls of aggressive ri
How SRE Teams Use Deliberate Practice to Sharpen Incident Response
How SRE Teams Use Deliberate Practice to Sharpen Incident Response Aug 15, 2026 8:16 In this episode, Lucas and Luna explore how site reliability engineering teams are borrowing a concept from music and sports: deliberate practice. They move beyond traditional game days and chaos engineering to discuss structured, repeatable training that builds muscle memory for incident response. The conversation centers on how a major cloud provider reduced its mean time to recovery by 30 perce
How SRE Teams Use Game Days to Stress-Test Incident Response
How SRE Teams Use Game Days to Stress-Test Incident Response Aug 15, 2026 9:44 In this episode of The Site Reliability Podcast, Lucas and Luna explore how SRE teams use game days—controlled, simulated incidents—to stress-test their incident response, decision-making, and coordination under pressure. They walk through a concrete example from a major payments platform that ran a quarterly game day simulating a database failover during peak traffic, revealing bottlenecks in com
How SRE Teams Use Chaos Engineering to Test Their Own Systems
How SRE Teams Use Chaos Engineering to Test Their Own Systems Aug 13, 2026 10:29 In this episode of The Site Reliability Podcast, Lucas and Luna explore chaos engineering as a disciplined practice for testing system resilience. They discuss the difference between chaos engineering and game days, walk through a concrete example of a chaos experiment on a payment gateway, and dig into how teams choose what to break, how to measure success, and how to avoid the chaos trap. They a

Recommended