The Art of SLOs with Alex Bramley

Google Cloud Platform Podcast - A podcast by Google Cloud Platform

Categories:

Today on the podcast, Jon Foust is back with Mark Mirchandani as we talk about SLOs and the importance of measuring service reliability with Alex Bramley. As a member of the Google SRE team, Alex and his coworkers help customers optimally run their services on Google Cloud. They collaborate with the client, weighing client needs and user needs to develop a plan that is affordable, efficient, and has the highest reliability for the user. Recently, they’ve been working to automate functions such as detection of outages, so that Google and the customer can work together quickly to get everything working smoothly again. Later, Alex, describes the steps developers go through at his workshop, The Art of SLOs, which was designed to help companies measure and improve reliability. At this workshop, attendees are encouraged to set SLO targets and error budgets. They are given theoretical reliability problems to solve, allowing them to practice without the added pressure of messy, real-world problems. The Art of SLOs helps developers understand what measurements are beneficial and why and the best way to implement projects that can take those measurements accurately. Alex was able to make the materials for the workshop free online! Alex Bramley Alex Bramley joined Google in January 2010 as the first Mobile SRE in London, after IBM bought the startup he enjoyed working for and made it much less fun. He spent around 7½ years in various reincarnations of Mobile/Android/Play SRE, looking after the infrastructure that makes phones smart, keeps them up to date, and provides them with countless distracting apps. CRE offered an interesting opportunity to do something different and learn from a bunch of very smart senior people, and Alex has not regretted taking the leap into the unknown. Much of his time recently has been spent rethinking how people teach customers, partners and the general public about SLOs. He helped create the Coursera course on measuring and managing reliability and developed what became the Art of SLOs for Liz Fong-Jones to deliver with other Google SREs at SREcon EMEA’18. Alex works four days a week so he can (suffer) enjoy looking after his children on Wednesdays, listen to cheerful music, and waste a lot of time playing computer games and occasionally writing code. Cool things of the week Postponing Google Cloud Next ’20: Digital Connect blog New research: How effective is basic account hygiene at preventing hijacking blog Simplified global game management: Introducing Game Servers blog Interview The Art of SLOs site CRE Life Lessons blog Putting customers first with SLIs and SLOs blog Putting customers first with SLIs and SLOs (Part 2) blog Measuring and Managing Reliability course Site Reliability Engineering books Question of the week How do I get started with GCGS? docs Google for Games Developer Summit Keynote video Google for Games Developer Summit Playlists videos Where can you find us next? Jon will be working on an Open Match sample project for the developer community. Mark will be making more videos like Error Reporting and error logging - Stack Doctor.

Visit the podcast's native language site