Brent Chapman
Engineering Better Incident Management
I make incident management less stressful and more scalable. I work with top enterprises to design, develop, and level up their incident management programs.
More Preparation, Less Stress
Incidents are inevitable. The real question is, are you ready for the next one? And did you learn everything you could from the last one?
I work closely with your engineering teams to develop, maintain, and upgrade your incident management capabilities, thereby reducing the impact, duration, and frequency of incidents.
Applied Expertise
I’ve led incident management transformations at Slack and Google, and I’ve worked with engineering teams at leading companies such as Webflow, Axon, Atlassian, and Indeed to develop and strengthen their incident management capabilities.
Outages and other incidents are costly in many different ways, including lost sales, reduced productivity, damage to reputation, and decreased employee morale.Â
As an expert in site reliability engineering, I’ve helped many organizations prevent these issues. However, despite everyone’s best efforts, incidents are unavoidable; sometimes things still go wrong. That’s why I specialize in incident management, ensuring that companies can resolve these problems quickly and effectively when they arise, and then learn everything they can from each incident.
Incidents are inevitable. The key is to be prepared for them and then to learn from every incident, strengthening your resilience for the future.
Let’s explore how I can help your company develop and strengthen these essential incident management capabilities.Â
Success Stories
I was the lucky beneficiary of Brent’s work on Incident Management at Google. His leadership and direct effort resulted in nothing short of a full reformulation of this key competency of SRE. […] Through Brent’s work and training we were able to build structure and automation around this critical area of expertise, which has resulted in significant reductions in MTTR and improvements in the org’s ability to learn and grow from its service interruptions (e.g. through postmortems).
Marc Alvidrez
Senior Staff Engineer and Senior Manager, Site Reliability, Google
I learned SO much about incident management from Brent.
When he joined Slack our incident response was chaotic; all hands on deck, uncoordinated and scary. Brent rolled up his sleeves and quickly introduced us to world class incident response.
These days we have a top notch response; baked in process, clear roles and responsibilities, automation and not half as scary.
V Brennan
Senior Director of Engineering and Regional Lead, Slack
Profile
I am an expert in emergency management for IT services, specializing in guiding companies to prevent, prepare for, respond to, and learn from emergencies. I work from a strong background in IT infrastructure, site reliability engineering (SRE), and public safety emergency management.
Slack recruited me to lead incident response and incident management for their Engineering organization and the company as a whole. I designed and built Slack’s incident management capabilities; ensured smooth day-to-day incident management operations; helped the company learn from, prevent, and prepare for incidents; and shared Slack’s incident management story with our customers and the industry.
As a leader in Google’s legendary SRE organization, I convinced senior management of the need to strengthen and standardize the company’s incident management practices. I created the Incident Management at Google (IMAG) system, which is now used throughout the company, and helped refine the Postmortems at Google (PMAG) system, which enables the company to learn from incidents of all sizes.
I bring a unique perspective to my work in IT, with experience as a former air search and rescue pilot and incident commander, an emergency dispatcher and dispatch supervisor for major arts and music festivals, and a Community Emergency Response Team (CERT) member and instructor.
Throughout my career, I have designed, built, managed, and scaled IT infrastructure and teams for various organizations, from startups to major corporations such as Google, Apple, Slack, Salesforce, Atlassian, and Netflix. I have collaborated with numerous organizations in Silicon Valley and globally, as well as with various non-profit and government entities. Additionally, I am the co-author of the highly regarded O’Reilly book Building Internet Firewalls, a developer of widely used open-source software, and a popular speaker at conferences worldwide.
My rare combination of experience as an emergency manager, technology manager, people manager, software developer, network/systems engineer, and educator allows me to quickly and effectively dive in, assess a situation, and deliver results.
Bring effective Incident Management to your organization
Learn about Great Circle’s consulting and training offerings
Great Circle Associates, Inc.
www.greatcircle.com info@greatcircle.com International: +1 415 861 3588 USA Toll Free: 877 GRT CRCL
Holly Allen
Scott Nelson Windels