When declaring an incident becomes everyone’s favorite workaround

You see someone declare a Sev-2 and you wonder: wait, why is that even an incident? Nothing is down. Customers aren’t affected. But a manager needed to get their team’s problem to the top of another team’s priority queue, and the incident process was a reliable way to make it happen. That’s not really what the incident process is for, but it worked, so where’s the harm?

The problem is, once folks see that this works, it starts happening more often. A product manager declares an incident because the incident notification is the fastest way to get leadership attention on a problem that’s been stuck in the backlog for weeks. An account team declares one because they need engineering support for a big demo to a major prospect and the incident process is the easiest way to pull engineers out of their sprint work on short notice. An engineer declares one because it’s easier than navigating the formal exception process for the deployment freeze.

The harm is cumulative. When a growing fraction of your declared “incidents” aren’t real emergencies, the urgency signal degrades. When a genuine Sev-1 arrives, people respond with less urgency because they’ve been conditioned to expect another workaround. And the incentives compound: folks who game the system get their problems solved faster, which teaches everyone else that gaming is how to get things done. Each individual declaration is an understandable decision by someone who needs to get something done; it’s the aggregate that corrodes the process.

Every one of these non-emergency declarations still carries the full overhead of a real incident. Responders get pulled off their planned work. Someone drops whatever else they were doing to serve as incident commander. Stakeholders context-switch to follow along. When you’re running enough of these, your teams are spending a meaningful fraction of their time in emergency mode for things that aren’t really emergencies, and all the indirect costs of incidents (disrupted projects, context-switching, recovery time) accumulate just the same.

There’s an irony here: people are reaching for the incident process because it works; they’ve seen that it reliably delivers coordination, prioritization, and urgency on demand.

The instinctive response is wrong

When companies notice this pattern, the instinctive response is often to tighten the declaration criteria. They add gatekeeping: maybe you need manager approval to declare an incident, or there’s a pre-declaration checklist you have to complete first, or someone reviews whether the declaration was “warranted” after the fact. The intent is reasonable. The net effect is corrosive.

Gatekeeping incident declarations is counterproductive. Every speedbump you build also slows down real incidents. The person who hesitates to declare because they’re not sure the problem is “bad enough” is already a common failure mode in incident response. Adding a formal approval step or a post-hoc review of whether the declaration was justified makes that hesitation worse, not better.

You also miss what the gaming is telling you: people reaching for the incident process are telling you that your normal processes are falling short. If you only crack down on the gaming, you suppress the symptom without learning anything from it, and the underlying problems persist.

Fix the escape routes, not the escaping

Instead, look at what side effects people are trying to trigger when they declare questionable incidents, and make those capabilities available through other means.

If the easiest way to bypass the deployment freeze is to declare an incident, create a non-incident exception process for urgent changes. This doesn’t have to be complicated; a lightweight approval from a designated release manager, with a clear escalation path, covers most cases.

If the easiest way to get your problem moved up another team’s priority queue is to declare an incident, create a prioritization escalation path that doesn’t require an incident. A cross-team triage meeting, an explicit expedite-request mechanism, or even a dedicated Slack channel that the right people actually monitor can absorb most of the pressure. The bar doesn’t have to be as high as “declare an emergency”; it just has to be lower than “wait six weeks for the next planning cycle.”

If the easiest way to assemble a cross-functional team on short notice is through the incident process, create a lightweight coordination mechanism for non-incident situations. Some companies call these “swarms” or “tiger teams” or “coordination requests.” The name doesn’t matter; what matters is that people have a way to get the collaboration they need without borrowing the incident process to do it.

Repeatedly gaming the incident process to get resource prioritization or cross-functional coordination isn’t a series of one-off workarounds; it’s a symptom of a systemic problem that needs a systemic response. Google’s SRE organization built formal Code Yellow and Code Red mechanisms for exactly this: structured ways to rally resources and elevate priority when a problem is serious enough to demand cross-functional attention, but isn’t an incident.

The diagnostic question

Look at your last dozen or so incidents and ask, for each one: was this declared because there was an emergency, or because the incident process was the easier path to something the team needed?

You don’t need a formal audit. Just ask a few experienced incident commanders and on-call engineers; they already know which ones were real and which ones weren’t. Then talk to the folks who called for the questionable ones (in a blameless, fact-finding way, of course). They’ll tell you exactly what’s missing from the normal processes, if you’re willing to listen.

People gaming the incident process is just a symptom. The underlying problem is usually that normal processes are too rigid, too slow, or too unresponsive, and the incident process is the path of least resistance. Fix the underlying problem and the gaming stops, because there’s nothing left to game around. Your incident urgency signal recovers, your teams stop burning emergency-mode cycles on non-emergencies, and when a real Sev-1 hits, people respond like it matters.

And if you’re dealing with this, take a moment to appreciate what it says about your incident process: people are borrowing it because it works. The fix isn’t to make it stop working. It’s to make everything else work that well too.

Incidents start before the response does

Your company has probably invested significantly in what happens after an incident is identified: incident response tooling, trained incident commanders, communication protocols, on-call rotations. That investment matters. But what about the gap between when a problem starts and when anyone on your team knows about it?

During that gap, customer damage is accumulating. The problem is getting worse, the blast radius is expanding, and nobody on the team is doing anything about it because nobody knows yet.

You can’t eliminate this gap entirely, but you can shrink it. Four investments make the biggest difference.

Broaden your detection surface

Automated monitoring is the first and best line of defense, but it can only catch the failure modes someone thought to check for. Human detection isn’t a gap you can eliminate; it’s a permanent and valuable part of your detection capability.

This means your customer support team is part of your detection infrastructure, whether or not you’ve told them so. So is any part of your company that interacts with customers regularly: account execs, customer success managers, even your social media team. They talk to your customers every day and often see concerns emerge before engineering does. And don’t overlook your customers themselves, who won’t limit their reports to your “official” support channels. If all these folks don’t have clear, fast escalation paths to flag potential problems for engineering, you have a detection gap that no amount of monitoring investment will close.

If your company is a heavy user of its own product, the detection surface extends even further. When I led Slack’s incident management program, literally anyone in the company might notice a problem while using Slack internally. Not every company is in that position (it depends entirely on what the product is), but those who are should take advantage of it. Make sure everyone (all the way down to the part-time security guard covering the front desk on weekends) knows how to report problems they see.

And watch for indirect signals. One of Slack’s best harbingers of “something is broken, even if we don’t know what yet” was the page-view rate on our public status page. If it started surging upward, we knew that something was wrong, even if we weren’t getting any other clear signals yet, and we’d start investigating. It was like smelling a light waft of smoke, well before the smoke detectors and fire alarms go off. If you have a public status page, consider adding its traffic patterns to your monitoring. A sudden spike in visits is a low-cost early warning powered by the collective behavior of your user base.

Lower barriers to reporting

Most of these detection channels depend on someone raising a concern, and that only works if the barrier to doing so is low. At many companies, the only mechanism for raising an alarm is to declare an incident, which triggers a full coordinated response: pages go out, a channel is created, an incident commander is assigned, people drop what they’re doing.

That’s appropriate when you know you have a real problem. But if the only way to raise a concern is to trigger that entire response, people will hesitate, and rightfully so. Nobody wants to be the person who launched a full incident response over a hunch that turns out to be wrong. So they wait for more evidence, and the detection gap grows.

Think of it like calling emergency services. When you call 911 (or 999, 000, 112, or whatever your country’s emergency number is), you don’t have to know whether you need an ambulance, a fire engine, a hazmat team, or a bomb squad. You describe what you see, and a trained dispatcher determines how serious the situation is, what sort of response is warranted, and who to send.

Your incident detection should work the same way: make it easy for anyone to say “I think something might be wrong,” and let someone with training, experience, and context determine what response is warranted. At Slack, introducing a lightweight mechanism for exactly this was one of the most impactful things we did.

Continuously right-size your alerting

It’s tempting to close the detection gap by making your monitoring more aggressive: lower the thresholds, add more alerts, page on anything that twitches. This can backfire badly. Every alert that wakes someone at 3 AM and turns out to be nothing makes it a little more tempting for your on-call engineers to dismiss the next one. Alert fatigue is one of the most insidious threats to detection, precisely because it accumulates gradually. Your alerting system doesn’t fail all at once; it erodes, one false alarm at a time, until the real alerts get lost in the noise.

The discipline runs in both directions: yes, add monitoring when you discover gaps, but regularly prune alerts that aren’t earning their keep. If a service-owning team can’t get through a review of every alert they received in the past week in a reasonable portion of a weekly ops review meeting, they’re getting too many alerts.

Examine the gap

Another way to shrink the detection gap over time is to examine it after every incident. You’re never going to be able to fully automate detection, but it’s still an ideal worth pursuing. Three questions, asked consistently in every post-incident review, create a steady stream of improvements:

  • How long was the gap between when the problem started and when we detected it?
  • Could we have detected it sooner?
  • What monitoring would we need to add, or what threshold would we need to adjust, to catch this kind of problem faster next time?

The bottom line

Investing in detection is investing in the foundation of your entire incident management capability. You can have well-trained incident commanders, practiced responders, and polished communication protocols, but none of it matters until you know there’s a problem.

Respecting fatigue isn’t coddling

Is it coddling when an on-call engineer takes the next morning off to recover after handling a production incident at 3 a.m., or is it a smart company managing a reliability risk?

Here’s what that night actually looks like. The engineer gets paged at 3 a.m., then spends two hours diagnosing the problem, coordinating with fellow responders, and restoring service. By 5 a.m., the incident is resolved and they get back to bed, but it takes them a while to settle down and get back to sleep.

Four hours later, they’re at standup. That afternoon, they’re in a planning meeting. That night, they’re still primary on the pager.

This is the default at most companies. Nobody made a deliberate decision that it should work this way; it’s just what happens when there’s no explicit policy for post-incident recovery. And it carries more risk than most leaders realize.

Incident response is more fatiguing than regular work

Responding to an incident isn’t like a normal day of developing features and chasing bug reports. The cognitive demands are qualitatively different: rapid context-switching under time pressure, high-stakes decisions with incomplete information, coordinating across multiple people and systems, all while knowing that users are affected and stakeholders are watching. And there’s a physiological dimension that regular engineering work rarely triggers: adrenaline. Incident response activates the body’s stress response in a way that writing code or reviewing a design doesn’t. That heightened state feels productive in the moment, but it depletes reserves fast, and the crash afterward is steeper than the apparent effort would justify.

This, incidentally, is one of the reasons that training and practice matter so much. Responders who’ve rehearsed the process and trust the framework around them experience a less intense stress response when real incidents hit. Turning incident response into a “routine emergency” doesn’t just improve efficiency; it reduces the physiological toll.

An engineer who’s been actively responding for a few hours isn’t just tired in the way that a long day makes you tired. They’re measurably less effective at exactly the skills incident response demands: integrating new information, evaluating competing hypotheses, making decisions under ambiguity, and recognizing when a current approach isn’t working.

The degradation is predictable

Responder fatigue follows a recognizable pattern. As it sets in, people stop processing new information as effectively. They agonize over decisions they’d normally make quickly, or they stop making decisions altogether. They develop tunnel vision, fixating on the one theory they’re already pursuing instead of stepping back to consider alternatives. They fall into ruts, essentially pursuing “Plan A, again, with more feeling this time” instead of asking whether Plan A is still the right plan. They get less creative, more rigid, and more prone to mistakes.

Fire departments study this, because it’s exactly the scenario firefighters face: interrupted sleep from overnight emergency calls, then back on duty the next day. The research consistently shows that the kind of fragmented, insufficient sleep they get around overnight calls degrades next-day cognitive performance to levels comparable to having had a couple of drinks. That’s why a growing number of fire departments have reconsidered their traditional 48-hour shifts; the performance degradation on day two is bad enough that departments are restructuring around it.

Self-reporting isn’t enough

The insidious part is that fatigue undermines exactly the capacity you need to recognize it. A fatigued responder genuinely believes they’re performing normally. That’s simply how fatigue works. “I’m fine, I can keep going” isn’t evidence of fitness. It’s one of the symptoms.

This is the critical organizational point. If a company leaves fatigue management to individual judgment (“take it easy if you need to”), it’s built a system that depends on impaired people accurately assessing their own impairment.

Aviation learned this the hard way. The FAA doesn’t ask pilots whether they feel too tired to fly. It sets hard limits on duty time and required rest periods, because decades of accident investigation proved that self-assessment under fatigue is unreliable. Pilots who’d been awake for 20 hours consistently reported feeling capable. The data said otherwise. If you’ve ever had a flight delayed while the airline sought a new crew because the original crew had “timed out,” you’ve seen these rules in action.

The tech industry hasn’t had its equivalent reckoning yet, but the same cognitive science applies. An engineer who handled a two-hour incident at 3 a.m. and says they’re fine at 9 a.m. may well believe it. That doesn’t mean they’re right, and building your next day around that assumption is a gamble most companies don’t realize they’re taking.

What active fatigue management looks like

Companies that take responder fatigue seriously don’t rely on individual heroism or self-assessment. They build a few specific practices into their incident management capability.

Explicit rest expectations. Not “take it easy if you need to,” but clear guidelines: an engineer who responds to a significant incident overnight is expected to start late or take the morning off, depending on duration and severity. The default is rest; working the next morning is the exception that requires a conscious choice, not the other way around.

The incident commander (IC) monitors for fatigue. During extended incidents, it’s the IC’s responsibility to watch for fatigue signals in responders: slowed decision-making, tunnel vision, repeated questions, irritability, loss of situational awareness. This is the same responsibility a fire officer has for monitoring crew fatigue on a fireground. A fatigued responder who stays on the line isn’t being dedicated; they’re becoming a risk to the response, their teammates, and themselves. On incidents that stretch beyond a few hours, this includes planning responder reliefs early rather than waiting for someone to admit they’re spent.

Promote the backup. After a significant overnight incident, consider moving the backup on-call engineer to primary for the next 12 to 24 hours. The person who spent two hours at 3 a.m. restoring service is not the person you want as your first line of defense if something else breaks that afternoon.

Rethink shift length. Most teams seem to default to week-long on-call shifts with several weeks between shifts, but there’s a strong case for shorter, more frequent shifts. The same logic driving fire departments away from 48-hour shifts applies: shorter shifts mean less accumulated fatigue per shift, even if each person’s total on-call hours per quarter are similar. Shift design is a fatigue management decision, whether your company treats it as one or not.

Some incident management platforms are starting to build this awareness into their tooling. incident.io, for example, detects overnight pages and proactively asks the responder the next day whether they’d like someone to cover their next shift. That’s the right instinct: making fatigue management a system-level concern rather than leaving it to the judgment of the person who’s least equipped to assess it.

It’s a reliability decision

Most companies aren’t actively choosing to ignore a fatigue problem. Rather, they have a fatigue problem that they haven’t noticed yet, because nobody has framed it as an operational risk. When a leader says “we trust our engineers to manage their own energy,” what they’re actually saying is: we have no organizational mechanism for ensuring that the people responding to our next incident are cognitively fit to do so.

Respecting fatigue isn’t coddling. It’s protecting the quality of everything your engineers do the next day, including the next incident response.

Heroic saves are near misses

It’s 3 a.m. in California, where most of the dev team are still snug in their beds. The auth system has started rejecting valid credentials. Early bird East Coast customers are already trying (and failing) to log in for the day, and thousands of users in Europe have already given up and gone elsewhere. In a couple of hours, the West Coast will be waking up too. A brilliant engineer swoops in and saves the day. She has legendary debugging skills and a deep understanding of the auth system, and she puts together a fix in forty minutes that would have taken anyone else hours to even diagnose. Later that morning, leadership is sending thank-you messages in the all-hands channel. Her VP awards her a small spot bonus, and her manager reminds her to include it in the next performance review cycle.

What doesn’t usually happen is anyone asking: what if she hadn’t been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.

Near misses look like successes

In aviation and other safety-critical fields, it’s widely accepted that a near miss is an unparalleled opportunity to learn and deserves the same investigation as an actual failure. The reasoning is straightforward: a near miss reveals the same systemic vulnerabilities that a failure does. The only difference between a near miss and a disaster is that the outcome happened to be good this time, often because of luck, timing, or the presence of one specific person.

A heroic incident response is a similar opportunity. The system nearly failed, and would have failed if that one engineer hadn’t been available or hadn’t known exactly what to do. Her skill, expertise, and dedication are worth appreciating. But her unavailability would have meant a much worse outcome, and that’s worth examining too. Too many companies celebrate the save and stop there.

The incentive nobody designed

When a company celebrates a heroic save without examining why the heroics were necessary, it sends a message. The message isn’t intentional, but it’s clear nonetheless: what gets valued is the dramatic rescue, not the boring preparedness work that would have made the rescue unnecessary.

Over time, that message shapes behavior. The engineer who writes thorough runbook documentation, trains new team members on the auth system, and invests in monitoring improvements doesn’t get the same recognition as the one who swoops in at 3 a.m. and saves the day. Preparedness work is largely invisible in performance reviews. Heroic saves are memorable.

The result is a perverse incentive loop. Heroics get rewarded, preparedness doesn’t, and the company remains dependent on heroic saves because nobody is investing in the alternative. This isn’t because anyone explicitly decided that preparedness doesn’t matter. It’s because the reward system is quietly rewarding the wrong thing, and nobody has noticed because the heroes keep delivering results. Until they don’t.

In my experience, this is one of the most common patterns in companies that are struggling with incident management. They have talented, dedicated people who keep delivering heroic results, and because the results keep coming, nobody realizes there’s a growing structural problem underneath.

The hero as single point of failure

The incentive loop creates a second problem. The hero gradually becomes a bottleneck and a single point of failure. When that engineer is on vacation and the next auth system incident hits, the team might spend hours just figuring out what’s going wrong, let alone fixing it. When they eventually leave the company (as they likely will; heroes tend to burn out), the team discovers that critical knowledge walked out the door with them.

I see this pattern regularly in my consulting work. In a company’s most serious incidents, it keeps turning to the same handful of heroic engineers. Those engineers are talented and committed, and their involvement has genuinely saved the company from significant damage. Everyone involved with incidents knows who they are, and breathes a sigh of relief when they join an incident channel. But the company has never seriously examined what its response capability looks like without them. The term that often comes up to describe these people is “indispensable,” which is really another way of saying that the company’s incident response capability depends on specific individuals’ availability.

Why the problem stays hidden

The most insidious aspect of this pattern is that it’s invisible to leadership for as long as the heroes keep delivering. Companies at the earliest stages of incident management maturity often don’t realize they’re at risk. Leadership sees consistently good outcomes and assumes the company has strong incident response, when what they actually have is strong individuals (and a certain amount of good luck).

By the time the fragility surfaces, the gap between where the company thought it was and where it actually was can be startling.

Heroic is a growth stage, not a compliment

When I assess incident management capabilities for my consulting clients, one of the dimensions I evaluate is program maturity: where is this company on the growth path from ad hoc response to reliable organizational capability? The first stage on that path is called “Heroic.” It isn’t meant to be flattering. It means that incident response quality is a property of specific talented individuals rather than a property of the company. When those individuals are available, things go well. When they’re not, things go sideways.

Every company starts here. The question is whether they invest in growing past it, converting individual capability into organizational capability. That transition is what the rest of the maturity model describes, and it’s the core of what effective incident management programs are designed to do.

What to recognize instead

None of this means companies should stop recognizing heroic contributions when they happen. When someone saves the day at 3 a.m., thank them. But also investigate why the heroics were necessary, and invest in the answers. That’s a form of recognition too: it says the save mattered enough to learn from.

To move from “Heroic” to higher levels of organizational capability, you need to shift what gets sustained recognition. Recognize the work that makes heroic saves unnecessary: the runbooks, the training, the well-coordinated responses where nobody had to be heroic.

If an engineer spent much of their quarter writing runbooks, training new responders, and coordinating incident responses, recognize that work: in performance reviews, in public acknowledgment from leadership, in awards and bonuses. If you don’t, you’re telling your organization that the only incident management work worth noticing is the dramatic save.

The goal is to make effective incident response something the company can do reliably, regardless of who happens to be on call. Heroes are still welcome, and still admired. They just shouldn’t be required.

Without a program to support them, incident management processes wither

Every fire department has a training program. Big-city departments have entire training divisions; even small volunteer departments that can’t spare anyone full time still name a training officer. Not because training is the department’s mission, but because maintaining the capability to do the mission requires sustained, dedicated attention.

New recruits need to be brought up to speed. Everyone needs to learn about evolving techniques and new equipment. Procedures need to be updated as building codes and materials change. Hard-won lessons from past incidents would survive only as stories told around the kitchen table; the fire service has a strong storytelling tradition, and its legends and cautionary tales carry real value, but oral history is hard to study, standardize, and train on.

Maintaining operational capability is itself a job, distinct from the operational work it supports, and fire departments size the role to the department rather than leave it unassigned.

Many software companies haven’t learned this yet. They invest real effort in building an incident management process. They define severity levels, write runbooks, designate incident commanders (ICs), set up communication channels. The project might take weeks or months of focused work, often driven by someone who cares deeply about doing it right (and often done in their “spare time”). When it’s done, it works, at least for a while. Incidents get declared. ICs run the response. Post-incident reviews happen. Everyone takes it for granted.

Then the person driving it gets promoted, or moves to another team, or leaves the company. The process, which was never really institutionalized because it didn’t need to be while that person was carrying it, begins to decay. Not catastrophically, but more like a garden nobody is tending any more: it doesn’t collapse overnight, it just slowly fills with weeds until one day you look up and realize the original design is barely recognizable.

The training materials haven’t been updated since the initial rollout. New engineers join but never go through incident training because nobody is scheduling it anymore. The severity level definitions still describe one product, but the company now has three. The IC rotation is running on the same six people it started with, even though the engineering team has doubled in size. The post-incident review template still references a tool the company stopped using a year ago.

None of these are crises on their own. Each one is easy to defer. But they compound, and the cumulative effect is that the process on paper bears less and less resemblance to what actually happens during incidents. In my experience, six months is roughly how long institutional momentum carries before the absence of active stewardship becomes visible in the quality of your incident responses. And growth accelerates the decay: the company simply grows away from the process, and nobody’s job is to notice.

This is what happens when you have a process but not a program.

A process is not a program

A process is a set of documented procedures: how incidents get declared, who fills which roles, what communication channels to use, how to run a post-incident review. A process can be written down, trained once, and followed.

A program is the organizational structure that develops, maintains, evolves, and champions the process over time. It’s the thing that keeps the process alive.

Many companies build the process and assume they’ve built the program. They haven’t. They’ve written a document, and documents don’t train new hires, don’t recruit for on-call rotations, and don’t update themselves when the company reorganizes around them. People do those things, and it only happens reliably when it’s actually somebody’s job.

“Everybody owns it” means nobody owns it

When I ask companies who owns their incident management program, the most common answer is some version of “we all do” or “the engineering organization as a whole.” This sounds collaborative. In practice, it means nobody has the explicit responsibility, the dedicated time, or the institutional authority to keep the process alive.

This organizational challenge isn’t unique to incident management. Companies that are serious about security don’t say “everybody owns security” and leave it at that. They assign ownership because shared responsibility without explicit ownership means the work doesn’t get done.

Incident management is the same kind of organizational capability. It needs someone whose actual job, not just their passionate side interest, is keeping it healthy.

What a program actually does

When I talk about an incident management program, I mean ownership of the full lifecycle of the capability, not just the procedures themselves. That includes keeping everything current as the company grows and changes: process documentation, severity definitions, escalation paths, tooling, runbooks.

It includes running a training pipeline so new hires are prepared before their first real incident, not thrown into the deep end during it. It includes maintaining the incident commander corps: recruiting new incident commanders, nurturing their development, supporting healthy on-call rotations across teams, and recognizing the people who do this demanding work. My former Slack colleague Scott Nelson Windels likens this to the farm teams and academies that elite sports clubs run: the point isn’t just fielding today’s roster, it’s making sure capable players are always coming up to fill it next quarter, too.

And it includes owning the post-incident review process and looking across incidents for patterns that no individual team would spot on their own. It includes tracking whether the process is actually being followed, and investigating when it isn’t, not to punish people, but to understand whether the process needs to change.

No single component is enough on its own, and no component stays healthy without sustained attention.

The good news

Building a program doesn’t require hiring a large team or creating a new department. At many companies, especially smaller ones, it starts with one person who has explicit ownership and dedicated time. What matters is that the responsibility is named, visible, and institutionally supported, not just assumed.

Here’s a quick test. Ask who owns your incident management program. Not who wrote the process, and not who ran the last big incident, but who is accountable, today, for whether the training is current, the rotations are staffed, and the severity levels still match the product. If the answer is a name, the follow-up question is what happens when that person leaves. If the answer is “everybody,” or someone who left the company last year, the process is quietly withering. And if you have a program but it would collapse without you, you haven’t finished building it yet.

The fire department didn’t name a training officer because it had extra budget. It named a training officer because it understood that maintaining a capability requires ongoing investment. The alternative, assuming trained firefighters stay trained and procedures stay current without anyone specifically owning those things, is how capabilities quietly erode until they fail when you need them most.


I’m writing a book on Incident Management for DevOps and SRE. If you’d like to know when it’s available, and get occasional updates along the way, you can sign up at im4ds.com.

If your company needs help building its incident management program, that’s the focus of my consulting practice at Great Circle.

Avoiding counterproductive pep talks

In the middle of a major incident, a senior leader joins the response channel and posts something like this:

“Let’s try to get this resolved in the next 10 minutes, please!”

They mean it as encouragement. Maybe they’re feeling pressure from their own leadership, or from a major customer, or both. Maybe they know the CEO is worried about a contract renewal call in an hour with a customer already frustrated with reliability. They want the team to know this incident matters. Posting a rallying message feels like leadership: it’s visible, it’s supportive in intent, and in normal day-to-day work, rallying the team really is a valuable leadership skill.

During an incident, it usually backfires.

What the responders actually hear

In nearly every incident I join, the responders are already working as fast as they feel they safely can. They don’t need to be told the incident matters; they’re the ones who got paged, who are staring at dashboards, who are juggling three competing hypotheses about what’s going on. When a leader urges that team to go faster, the message they receive isn’t “we believe in you.” It’s “you’re not working hard enough.”

And if the team is already at their limit, a call for more speed can’t add speed. It can only add anxiety. Anxiety during an incident is expensive: responders start splitting their attention between the problem in front of them and the audience watching them, and split attention is something that a complex technical investigation can’t afford. The pep talk was meant to help the team focus. It does the opposite.

The trouble with “10 minutes”

The arbitrary deadline makes it worse. Why 10 minutes? Where did that number come from? The responders don’t know. Often the leader who posted it doesn’t know either; it just sounded suitably urgent.

But now the number is sitting in the channel, and every responder is doing math against it. If the team resolves the incident in 12 minutes instead of 10, did they fail? Nobody can answer that, which means the deadline has created a test that everyone can feel and nobody can pass. Some part of each responder’s attention now goes to the clock, and to the question of how this will look to a senior leader who’ll have a big say in their next performance review. None of that attention is going to the outage anymore. The team has been handed a no-win condition in the middle of an emergency, by someone who was trying to help.

How experienced commanders convey urgency

I’ve spent a lot of time around public safety incident command, and one of the things that struck me early is what you don’t hear on the radio at a working fire: nobody broadcasts “let’s try to knock this fire down in the next 10 minutes, please!” to the crews working the fire.

What you do hear is information. “We have a report of a person trapped on the second floor” changes how the crews operate, instantly, without anyone being exhorted to care more. On the fireground, commanders convey urgency through facts and objectives, because facts and objectives change what responders do. Cheerleading doesn’t.

Even with the arbitrary number removed, “let’s wrap this up quickly, team!” is not actionable, because it doesn’t tell anyone what to do differently. It changes the mood, but probably not for the better, without changing a single decision.

Urgency is information, not exhortation

The leader’s urgency is usually genuine, and often there’s real business context behind it. That context is valuable. But it needs to be delivered as information, to the right person, through the right channel.

The right person is the incident commander (IC), the person coordinating the response. The right channel is a private one. “The CEO has a contract renewal call with our largest customer in an hour; they’re already frustrated with our reliability, and this outage isn’t going to help” is genuinely useful: the IC can prioritize mitigations that affect that customer’s services, loop in the account team, or prepare a status update the CEO can reference on the call. “Legal needs to know by end of day whether customer data was affected, so they can meet the notification deadlines in our customer contracts” is useful in the same way. The IC can act on information like that.

What to do instead

For senior leaders: when you feel that urge to rally the troops, pause and apply a simple test: can the responders use what you’re about to post to make any decision better? If so, you have real business context, and you already know where it goes: to the IC, privately. If not, it might indeed change the mood, but probably not for the better. And if the honest answer to what’s driving your urgency is “I’m anxious and I want them to know I’m paying attention,” then the most supportive thing you can do is trust the team and stay out of the channel. The most valuable contributions senior leaders make during major incidents mostly happen outside the response channel anyway: clearing roadblocks and handling the stakeholders who would otherwise be pestering the responders for updates.

For incident commanders: when one of these pep talks lands in your channel, don’t respond defensively, but don’t ignore it either. Acknowledge it briefly, then follow up with the leader privately: “Is there specific business context driving that timeline? If so, it would help me to know what it is.” Most of the time there is something behind it, and that question converts an anxiety-inducing exhortation into information you can actually use. You’ll also be quietly teaching your leadership how to engage with the next incident.

For everyone else on the response: you don’t need to respond to the pep talk, and it doesn’t change your priorities unless the IC says it does. During an incident, you take your cues from the IC, not from voices outside the response, no matter how senior. Messages like this are the IC’s to handle, and now you know how they’ll handle it.

The urgency itself was never the problem. Every incident deserves urgency. The problem is urgency delivered as pressure instead of as information, because pressure makes responders slower and more mistake-prone just when the company most needs them sharp.


I’m writing a book about all of this: Incident Management for DevOps and SRE. If you’d like to hear when it’s available, you can sign up at im4ds.com.

Need help preventing, preparing for, responding to, and learning from incidents? That’s the focus of my consulting practice at Great Circle.

Area Command: what to do when incidents collide

Most of the time, simultaneous incidents are no big deal. Any company with enough engineers and enough services will have multiple incidents open at the same time; when I led incident management at Slack, we had a few thousand engineers, and half a dozen concurrent incidents was not unusual for a typical weekday. Each had its own incident commander (IC), its own responders, its own channel, and they proceeded independently without anyone needing to think about the others.

Any conflicts that emerged over priority or resources could usually be worked out among the incident commanders of the individual incidents. The most obvious approaches are straightforward: a higher-severity incident takes priority over a lower-severity one, or a customer-impacting incident takes priority over an internal-only one. When only two incidents are in contention, the two ICs can almost always negotiate a decision between themselves.

Sometimes, though, negotiation among the ICs couldn’t produce a timely solution, particularly when more than two incidents were involved.

Consider a large SaaS platform with three active incidents, all with fixes ready, all needing the same deploy pipeline. Incident A is a Sev-1: the customer-facing API is returning errors for a segment of users. Incident B is also a Sev-1: a partner-facing integration is down. Incident C is a Sev-2: a background data pipeline is falling further behind, and if it isn’t addressed in the next few hours, the backlog will cascade into customer-impacting failures worse than A or B. The severity heuristic doesn’t break the tie between A and B. The customer-impact heuristic doesn’t help either: A and B are both customer-impacting in different ways, and C will be soon. With two incidents, the two ICs can talk it through. With three, each IC is advocating for their own incident, and the tradeoffs involve factors that no individual IC has visibility into: the contractual implications of A, the partner relationship at stake in B, the rate at which C’s backlog is growing. Somebody with a cross-incident perspective needs to decide the sequencing.

When I was at Slack, this kind of contention came up about twice a year. The trigger was usually the deploy pipeline for the monolith. Normally, incidents deployed fixes alongside each other and routine code pushes without any issues. But occasionally, an incident’s fix was risky enough to need exclusive use of the pipeline: a staged rollout that could take a couple of hours, where we didn’t want to bypass the staging except in a dire emergency, and didn’t want to bundle high-risk fixes for multiple incidents into the same deploy because rolling back one fix mid-deploy could derail the other. When multiple incidents each wanted exclusive use of the pipeline at the same time, the ICs were stuck.

Three options when incidents collide

When simultaneous incidents start interfering with each other, there are three options.

Leave them separate. If the interference is minor and manageable (one incident can wait an hour for the deploy pipeline, or the resource contention resolves with a quick conversation between the ICs), there may be nothing to do beyond agreeing on a path forward. Not every collision warrants escalation.

Combine them. When investigation reveals that two “separate” incidents share a common underlying cause, consider merging them into a single response. What looked like independent problems turns out to be different symptoms of the same failure, and maintaining separate responses would mean duplicating effort and coordination. But don’t be in too big a hurry to merge: concurrent incidents are common, and correlated incidents are much less common. Merging is also hard to undo; if you realize mid-response that the incidents weren’t actually related, unmerging is about as messy as reopening a closed incident (pro tip: don’t try, open a new incident instead).

Stand up a coordination layer. When multiple active incidents are genuinely distinct (different systems, different expertise needed), but they’re competing for the same scarce resources, and the individual ICs can’t resolve the contention among themselves, somebody needs to make the cross-incident prioritization calls. Keep the incidents separate but add a coordinator above them to make priority, policy, and resource allocation decisions. In the Incident Command System, this coordination layer is called Area Command.

Area Command

To see why the Area Command pattern is needed, consider a tornado outbreak that drops multiple tornadoes across a metro area in quick succession. A school gym hosting a basketball game has partially collapsed with people trapped inside. A residential neighborhood has been leveled, with reports of injuries throughout. A gas line rupture in an industrial area has started a fire that’s threatening an adjacent apartment complex. Each of these is a separate incident with its own IC and responders, but there aren’t enough responders to go around; all three incidents need the same scarce resources.

Heavy-rescue teams are needed at the school gym to reach trapped spectators, but also in the residential neighborhood where people are buried in collapsed houses. Fire engines are needed at the gas fire to keep it from reaching the apartment complex, but the residential neighborhood also has secondary fires breaking out. All three incidents need ambulances. None of the individual incident commanders is in a position to make the tradeoffs across incidents; each one is rightly focused on their own scene. Somebody above them needs to make the hard calls: which scene gets the heavy-rescue teams first? How many of the fire engines go to the industrial fire and how many get sent to the neighborhood? That’s Area Command.

The structure is deliberately lean. The Area Commander doesn’t manage any incident directly. They allocate scarce resources across incidents, set relative priorities, and communicate the aggregate picture to elected officials. Each incident still has its own IC running its own response; Area Command coordinates between them, not within them.

What this looked like at Slack

At Slack, we adapted Area Command for situations where multiple incidents collided and needed coordination above the level of the individual ICs. The Area Command was itself an incident, with its own incident number, its own channel, and its own IC (the “Area Commander,” typically myself or another senior member of the incident management program). The Area Commander would interface with the IC from each active incident (or the IC would name a liaison) and focus on the questions that no individual IC could answer alone: which incident should get priority for deploying fixes, how to consolidate conversations with executive leadership, and whether a scarce resource should stay on one incident or move to another.

These are sacrifice decisions at the company-wide level: deliberately accepting a worse outcome on one incident to get a better outcome on another. The Area Commander would frame the decision for executives when needed, then work with the individual incidents to implement whatever was decided.

What made this work was keeping it lean and temporary. There was no Tech Lead at the Area Command level, because the technical work was happening within the individual incidents, not at the coordination level. If an IC designated a liaison to Area Command rather than handling that themselves, it was never the Tech Lead. The TL’s attention stayed on their incident’s technical workstreams.

We activated Area Command roughly twice a year, typically only for about an hour. It spun up when incidents started tripping over each other, and spun down as soon as the contention was resolved (often well before the incidents themselves were resolved). It’s not a standing organization; it’s a coordination pattern you reach for when you need it and put away when you don’t.

A rare but high-stakes situation

Most companies won’t need Area Command often. But the situations where it’s needed are precisely the situations where you don’t want to be figuring it out from scratch.

The pattern itself is simple: keep the individual incidents running independently, stand up a lightweight coordinator above them, and give that coordinator the authority and information to make the tradeoff calls that no individual IC can make. It doesn’t require elaborate tooling or training. It requires someone senior enough to make prioritization and resource allocation decisions, a channel for the coordination conversation, and a line of communication with each active incident.

The harder part is recognizing when simultaneous incidents have crossed from “running at the same time” into “interfering with each other.” Most companies stay in the first category most of the time. But any company with enough services and enough engineers will eventually hit the second, and having a name for the coordination pattern (and having thought through how it would work) makes the difference between a structured response and several ICs each lobbying for their own incident with no one positioned to make the call.


Area Command is one of the patterns I cover in my forthcoming book, “Incident Management for DevOps and SRE.” If you want to know when it’s available, sign up at im4ds.com.

If your company needs help building or improving your incident management practice, that’s the focus of my consulting at Great Circle.

Workstreams, for when your incident channel gets too congested

Most companies’ incident processes work well for routine incidents. An engineer gets paged, a few colleagues join the response, someone takes on the incident commander (IC) role, and the team works the problem together in a Slack channel. The IC can keep track of what everyone is doing because “everyone” is five or six active responders. Communication flows naturally. The process feels lightweight because it is.

The first scaling step: naming a tech lead

The first time this setup gets strained is usually when the IC has too many things competing for their attention. They’re trying to guide the technical investigation while also fielding questions from stakeholders, coordinating communication, and keeping the response organized. The IC’s attention fragments, and both the technical work and the coordination suffer.

The answer at this stage is naming a tech lead (TL) for the incident: someone who takes over direct management of the technical work while the IC handles everything outward-facing. Together, an IC and TL can effectively coordinate more responders than either could alone, because they’ve split the two biggest demands on attention into separate roles. Adding a TL helps you scale beyond what the IC alone could handle, but it only takes you so far.

When even IC and TL aren’t enough

When an incident has enough active responders working on different aspects of the problem simultaneously, even an IC and TL working together can’t directly coordinate all of them. The single channel becomes congested with interleaved conversations about different problem areas. Responders spend more time trying to follow the scroll than working the problem. People start working at cross purposes because nobody has visibility into what other responders are doing.

To be clear: it’s fine to have dozens of people watching an incident channel. Spectators following along for situational awareness is a sign of healthy transparency, not a coordination problem. The challenge is the active responders, the people actually working the incident, and what happens when there are more of them than an IC and TL can directly coordinate.

What experienced responders do next

In many companies, the experienced responders sense this breakdown before anyone names it. They’ve seen it before. They get frustrated with the chaos, quietly break off into smaller groups, and start working the problem in DMs, side huddles, or breakout channels where they can actually focus.

What they’re doing, whether they’d use the word or not, is forming workstreams: small teams organized around specific aspects of the problem. For example, in an incident involving data corruption across multiple systems, one workstream might focus on stopping the corruption, another on assessing the blast radius, and a third on coordinating customer communication. Each workstream has a natural focus and can work without wading through the noise of the other workstreams’ conversations.

This instinct is sound. The main channel has become too congested for focused technical work, and breaking into smaller, objective-focused workstream groups is exactly the right response. But when it happens informally, it creates its own problems. The IC and TL may not know the workstreams have formed. There’s no explicit coordination between the workstreams. Information that matters to one workstream gets trapped in another workstream’s side conversation. And at its worst, it shades into freelancing: experienced people working on what they think is most important, outside the IC’s and TL’s awareness, with nobody having the big picture across the effort. The whole thing only works when those particular experienced people happen to be responding. When they’re not, the incident stays in the chaotic single-channel mode, and the response suffers for it.

Making the instinct explicit

The fix is to formalize what experienced responders already do naturally. When an incident grows beyond what an IC and TL can directly coordinate (roughly eight to ten active responders, in my experience), break the response into named workstreams, each with a designated workstream lead who coordinates the work within that workstream.

The TL’s role shifts at this point. Instead of directly managing all the technical responders, the TL coordinates across workstream leads: connecting dots between workstreams, making sure one workstream’s fix isn’t creating problems for another, and maintaining the technical big picture that no single workstream can see. The IC continues to handle organizational coordination, stakeholder communication, and the overall direction of the response. It’s the same IC/TL partnership, scaled up one level.

One detail matters more than it might seem: name workstreams after what they’re trying to accomplish, not which team the people came from. “Stop the data corruption” is a workstream objective. “Database team” is an org chart label. In a complex incident, the work rarely falls along team boundaries; the whole point of the incident response program is to be able to form an ad hoc team to address the problems. A workstream focused on stopping data corruption might need engineers from the database team, the networking team, and the application team all working together. Naming by objective keeps everyone oriented toward the same goal and avoids the trap of siloing along reporting lines when the problem doesn’t respect those lines.

How fire departments scale their coordination

Fire departments face the same scaling challenge, and the way they handle it is instructive. The Incident Command System (ICS) explicitly defines how the organizational structure scales with the size of the incident. A single-engine response to a dumpster fire has one officer managing a 3-4 person crew. A multi-alarm apartment building fire has an IC at the top of an org chart spanning multiple engine and truck companies, rescue teams, ambulances, and other specialized units organized across multiple floors and faces of the building. As units arrive to join the response, the organizational structure scales up in well-defined ways to absorb them, and firefighters train for these transitions before they ever face a live fire.

The specific structure from ICS doesn’t translate directly to software incidents (we don’t need strike team leaders, for example). But the principle does: the coordination model must evolve in predictable, preplanned ways, and people need to practice making those changes before they’re in the middle of a crisis.

Getting started with workstreams

Most companies design their incident process for the incidents they have most often, and that’s reasonable. Routine incidents with a handful of responders don’t need workstream coordination. But if you’ve ever had an incident where the response felt scattered and chaotic, where experienced people quietly broke away to work independently, and where nobody had the big picture across all the parallel efforts (or even knew what all those efforts were), the issue was probably that your coordination model didn’t scale with the incident.

The shift to workstreams doesn’t happen all at once in an incident. When the IC or TL recognizes that the response has outgrown direct coordination, they might start by spinning up a single workstream for the most clearly defined problem area, while other responders continue working in the main channel under the TL’s direct coordination. As more distinct problem areas emerge, more workstreams form. It’s a gradual transition, not an instant cutover. And the reverse applies as well: as workstreams accomplish their objectives, they can be dissolved and their responders absorbed into the remaining workstreams or stood down. The coordination structure should contract as the incident winds down, just as it expanded when the incident grew.

A few practical starting points: when you spin up a workstream, give it its own channel so the work is visible, not DMs or side huddles where it disappears from the IC’s and TL’s view. Name it by its objective. Each workstream should have an identified lead who coordinates the work within the workstream and reports back to the TL. The TL connects the dots and coordinates across workstreams, making sure one workstream’s approach isn’t undermining another’s. And just like with Slack threads, key discoveries and decisions made within a workstream need to get shared back to the main channel immediately, so the IC, the TL, and other workstreams all have the full picture.

There’s more to managing workstreams well than a blog post can cover: how workstream leads communicate status, how the TL balances attention across workstreams, how to handle responders who need to move between workstreams as the situation evolves. I cover these in a chapter on large, long-running, and other special-situation incidents in my forthcoming book, Incident Management for DevOps and SRE. Sign up for publication updates at im4ds.com. If your company needs help building this into your incident process right now, my consulting practice is greatcircle.com/im.

The experienced responders on your team already instinctively self-organize when an incident gets big. Your process should be supporting and channeling that instinct into something reliable and repeatable, not leaving it to chance.

Modern software architecture means nobody has the whole picture. To assemble one in an emergency, you need an incident tech lead.

Every large software development organization has made the same bargain: if nobody has to understand the whole system, we can build a more capable system than we otherwise could, even though it grows bigger and more complex. We break systems into components with well-defined interfaces so that each team needs to understand only the pieces it owns, plus the interfaces of its neighbors. That’s the point of every decomposition strategy, whether it’s microservices, bounded contexts, service ownership, or a well-modularized monolith. The architecture deliberately limits what any one person has to hold in their head.

It’s a sound strategy. It’s also why some incidents are so much harder than others.

Lorin Hochstein named this pattern beautifully in a recent post, The demon of the gaps. Failures that stay inside a single component are the easy ones; you page the owning team, they figure out what’s wrong, and they fix it. The hairy incidents emerge from unexpected interactions across components: several services throwing errors at once, or no services throwing errors while customers see broken behavior anyway. As Lorin puts it, “you’ve built an analysis solution but you’re now faced with a synthesis problem.” In order to scale, the architecture deliberately optimized away the need for whole-system understanding; now the whole system isn’t working, but nobody has that understanding to call on.

I’ve watched this play out in incident channels many times. Subject matter experts from six different teams, each reporting that their own service looks healthy. Six dashboards are green, but meanwhile, checkout is still failing for customers. The knowledge needed to explain what’s happening exists, distributed across six heads, but nobody is assembling the pieces. The responders need to understand how the system as a whole is behaving right now. That understanding has to be built live, under pressure, from multiple partial models. That’s synthesis work, and it doesn’t happen on its own.

Lorin observes that guidance on preparing for this work is almost nonexistent. Here’s the encouraging part: closing the structural gap is fairly straightforward. Most companies already use structured incident roles (incident commander, subject matter expert, customer liaison, etc.); they need to add a synthesis role, activated when needed for complex incidents.

Synthesis is a job. Name it.

The incident commander (IC) coordinates the overall response; every incident will have one. On the most complex incidents, though, where you need this synthesis function most, the trick is to also activate an incident tech lead (TL) to lead the technical investigation. The role is analogous to the tech lead role many teams have in their everyday structure, but its scope is the incident rather than one particular service. Most companies have never established the incident TL role, and for routine incidents they don’t miss it: the IC can handle the technical side along with everything else, but for complex incidents, the TL role can be incredibly valuable.

The incident TL job, properly understood, is the synthesis job: connecting observations across component boundaries, correlating the partial models from different subject matter experts, and maintaining the evolving picture of how the system is failing and what we’re doing about it. The TL doesn’t need to be the deepest expert in any single component. They need to be good at building a working model out of other people’s expertise, and much of that work is cross-checking, holding indications from different components up against each other and noticing the discrepancies: “If we’re seeing this in component A, we should be seeing that in component B, but we aren’t; why not?” “If A is doing this and B is doing that, the problem must be upstream of both.” “Wait, A says one thing but B says another; they can’t both be right, can they?”

The separation between IC and TL exists to protect that work, and it cuts both ways. Synthesis requires sustained, heads-down attention; you can’t reconstruct a system model in the gaps between stakeholder updates and staffing decisions. And the same complexity that makes an incident demand serious synthesis also multiplies the outward-facing work: more stakeholders to update, more escalations, more decisions about the response itself. The two loads peak together, and one person can’t carry both.

The IC takes everything outward-facing precisely so the TL can stay immersed in the technical picture, and the TL handles the heads-down focused work so that the IC has time for everything else. When I’m the incident commander, one of the most valuable things I can do for my tech lead is keep everyone else out of their hair. But the separation is a division of labor, not a wall. I like to think of the IC and the TL standing back to back, facing opposite directions, talking over their shoulders to keep each other informed. Each is watching a different part of the horizon, and together they have the whole picture.

The response team crosses the boundaries on purpose

An incident response is a temporary organization: an ad hoc team assembled across ownership boundaries for exactly as long as the incident lasts. Conway’s law observes that systems end up mirroring the communication structures of the organizations that build them, and the mirror works in both directions: your team boundaries and your component boundaries align, which is exactly what you want for everyday work. The incident structure deliberately cuts across those boundaries, because the gaps between components are where the problem lives. Pulling six SMEs into one channel isn’t enough by itself, though. A group of experts in the same room is a meeting; a group of experts with someone responsible for synthesizing what they know is a response.

The communication mechanisms are synthesis tools

The standard incident communication practices may seem like bureaucratic overhead until you see what they’re for. “Going around the horn” (each responder, in turn, briefly reports what they’re seeing and doing) forces the partial models into the open, where the TL can correlate them. A periodic situation report, or SitRep, forces someone to compress the current understanding into a few sentences; writing it is itself an act of synthesis, and reading it gives every responder the same baseline picture to work from. Narrating before you act keeps each responder’s local view visible to the whole room. None of these mechanisms exists for discipline’s sake. They’re how a group of people, each holding a partial model, builds and maintains a shared one.

Wildfires don’t respect organizational boundaries either

As is often the case in incident management, we can look to fire departments for inspiration and solutions. Consider a major wildfire. Dozens of agencies converge: federal, state, tribal, and local, some from hundreds or even thousands of miles away. No single agency understands the whole incident, with its terrain, weather, fuel, crews, and aircraft. The Incident Command System (ICS), the standard structure for emergency response in the US and beyond, is how all these disparate parts get pulled together into a coherent whole. ICS treats building the shared picture as a staffed function: a planning section tracks the situation and the resources, assembles the common operating picture, and distributes it to every responder through the incident action plan. Nobody simply hopes that shared understanding will emerge; somebody owns producing it.

Software companies can borrow that lesson directly: treat synthesis as a named responsibility rather than an emergent property. If the IC role at your company is defined as “project manager of the outage” and nobody is explicitly responsible for assembling the technical picture, the synthesis function is unowned, and it will show in your cross-boundary incidents. Establish the incident tech lead role. Protect it from outward-facing distraction. And practice it: when you run game days or tabletop exercises, choose scenarios that cross team boundaries, because those are the scenarios that exercise synthesis rather than component expertise.

Decomposition made whole-system understanding nobody’s everyday job, and that’s fine; it’s a good strategy with a known cost. Incident management structure is how you pay that cost only when you must, with machinery built for the moment.


I’m writing a book, “Incident Management for DevOps and SRE.” Sign up at im4ds.com to be notified when it’s available, and to get occasional progress updates and early access to selected content.

If your company needs help with incident management right now, that’s the focus of my consulting practice at GreatCircle.com/im.

The on-call cost of AI-generated code

If your engineers are using AI coding assistants, your team is almost certainly shipping more code than they were before adopting these tools. That’s not surprising: the whole point of these tools is to accelerate how fast code moves from idea to production. The velocity story is real, and it’s the story most companies focus on.

The key question is, when that new code breaks in production at 3am, how well can the on-call engineers debug it?

The understanding gap

I’ve written before about how AI tools are quietly thinning the understanding that teams have of their own systems. The short version: AI-assisted development shifts how code gets produced in ways that leave the team with shallower collective knowledge of the codebase. Not because anyone is doing something wrong. Good teams still do design reviews, still do code review, still write documentation.

But when AI generates code, the team reviews the output rather than participating in the implementation choices. The understanding they build is real, but it’s not as deep as what they’d have if they’d built it together. TR Jordan of Tern captures the shift well: the old deal was that if it was worth your time to write the code, it was worth my time to read it. When the code is AI-generated, there’s so much more code to review that the deal breaks down, and the knowledge-sharing that used to be baked into the process has to be rebuilt deliberately.

During normal operations, that’s fine. Teams have time to read through unfamiliar code, query the AI, run experiments, consult documentation. The pace is forgiving.

When thinner understanding meets time pressure

The pager goes off at 3am, and within minutes the response becomes a team effort: the on-call engineer pulls in teammates, the incident tech lead drives the investigation, subject matter experts get paged. But the team’s effectiveness under pressure depends on their collective understanding of the systems and code involved. That understanding is exactly what’s gotten thinner as the code volume has increased and more of the codebase has been shaped by AI.

This doesn’t mean the team is helpless. They can still read the code, still query the AI about what it does, still use their debugging tools. But there’s a difference between understanding code well enough to work with it during the normal course of development and understanding it well enough to reason about its failure modes at 3am, under time pressure, with customers affected. The first is a comfortable margin. The second is where gaps in understanding become visible.

The more of the codebase that’s been shaped by AI, the more the incident response team is working in territory they know less deeply than they would have if they’d built it all themselves. Each individual piece of AI-generated code might be fine. But in aggregate, the team’s ratio of “code in production” to “code we understand deeply enough to debug under pressure” has shifted. And it’s shifted in the wrong direction for incident response.

From valuable to essential

Firefighters deal with a version of this problem every time they respond to a fire in a building they’ve never been inside. They don’t know the floor plan, the hazards, or the building’s history. What they rely on instead are general diagnostic skills: understanding building types and construction methods, knowing how fire behaves, reading smoke conditions and other indicators. They’ve trained specifically for navigating the unfamiliar, because in their line of work, the unfamiliar is the norm. And they don’t just rely on those skills in the moment. Between calls, they prepare: conducting familiarization visits to buildings in their district, having informal “what if?” discussions over the kitchen table, running whiteboard sessions, reviewing and updating pre-incident plans. They build as much understanding as they can before the alarm sounds, knowing it won’t be complete but also knowing that every bit of preparation helps.

The Google SRE book describes an analogous training approach for software engineers: building the general skill of dropping into an unfamiliar system under pressure. Using diagnostic tools and debugging surfaces. Following requests across service boundaries. Drawing inferences from logs and metrics. Making that process reflexive enough to work when the stakes are high and the clock is running.

That skill set has always been valuable, but AI-assisted development makes it essential. When a growing share of your production code was written or substantially shaped by AI, the ability to debug systems you didn’t build is no longer just a nice-to-have that distinguished your strongest engineers; it’s a core competency your entire on-call team needs.

Of course, this assumes you’ve invested in the infrastructure to support those skills: diagnostic tooling, distributed tracing, structured logging, debugging surfaces that actually reveal what’s happening across service boundaries. If your company is shipping more AI-generated code, the case for investing in observability infrastructure gets stronger, not weaker. The skills and the tooling go together.

What this means for your company

If your company is adopting AI coding tools, the question isn’t whether the understanding gap exists. It’s whether your incident management practices account for it.

Invest in general diagnostic skills. Don’t just train engineers on specific systems; train them to navigate unfamiliar ones. Structured debugging exercises, shadowing across teams, and practice with diagnostic tooling all build the kind of transferable skill that matters most when the code is unfamiliar.

Don’t assume familiarity will come from the work itself. When teams hand-wrote most of their code, system understanding was a natural byproduct of the development process. AI-assisted development weakens that link. Companies need to explicitly invest in building the shared understanding that used to come for free. Some of that investment is formal: structured on-call ramp-up, cross-team shadowing, and light-weight training exercises. But some of it is informal, and just as important: engineers walking each other through recent changes, pairing on debugging sessions, having “what would we do if X broke?” conversations over lunch.

Build understanding between incidents. Firefighters build a lot of their knowledge around the kitchen table between calls. Software teams need the equivalent, and they need to protect the time for it. Dedicate a regular slot in your weekly team meetings for disaster role-playing or system walkthroughs. Google’s SRE teams have done this for years with a practice they call “Wheel of Misfortune”. The key is, it’s not a big-deal formal exercise, it’s just how they spend the last ten minutes of a weekly meeting.

Treat this as an organizational capability problem. Adopting AI coding tools for velocity gains is an organizational decision. So is investing in the operational readiness to match. That’s not an argument against AI tools; it’s an argument for thinking about the full picture. Shipping more and faster is valuable. But the cost shows up at 3am, when code your team doesn’t fully understand breaks in production and the clock starts running.


I’m writing a book on incident management for DevOps and SRE that covers this and much more. Sign up at im4ds.com to be notified when it’s available.

If your company needs help preventing, preparing for, responding to, and learning from incidents, my consulting practice is greatcircle.com/im.

AI ops tools are quietly eroding the awareness teams need during incidents

AI is automating operational work at an accelerating pace. AI ops tools handle monitoring, remediation, environment management, and infrastructure tasks that engineers used to do themselves. The “AI SRE” product category barely existed two years ago; now every vendor in the space has one. AI-assisted development tools generate code, suggest architectures, and handle implementation details. These tools deliver real value; they’re saving teams real time on real work today.

But there’s a second-order effect that most companies deploying these tools aren’t accounting for: the operational work that AI is taking over was also how engineers unconsciously built the system knowledge they need during incidents.

The engineer who regularly works with the infrastructure (designing, deploying, scaling, troubleshooting, tuning, investigating when things look wrong) develops an intimate knowledge of the systems: the dependency chains, the failure modes, what “healthy” looks like.

Every generation of automation has eroded some of that knowledge, and that’s usually been a worthwhile trade. Auto-scaling is a good example: it works so well that nobody thinks about scaling behavior day-to-day. Right up until the auto-scaler walks off a cliff, spawning so many new frontends that they overwhelm the database backend with connection requests and cache warmups, then time out and abort before getting online, wasting all the work the backend did to try to start them, and kicking off a crash loop of attempting to start, timing out, and retrying. The engineers responding to that incident need to understand scaling dynamics that haven’t been part of anyone’s daily awareness since the auto-scaler took over. That pattern predates AI entirely.

But AI is automating a broader range of operational work, faster, and the knowledge that erodes with it is correspondingly deeper.

Here’s why that matters for incidents: incidents are, by definition, the situations that the automation can’t handle. They’re what’s left over when everything that could be automated has been. And as the automation (both traditional and AI) gets more capable, the left-overs get messier and more complicated. The people who need to respond to those situations are the same people whose day-to-day work is increasingly mediated by AI. They have less deep understanding of the systems they’re being asked to debug, investigate, and reason about under pressure.

This is true even if you aren’t using any AI tools during incident response itself. The awareness erosion happened before the incident started.

Part of why this is happening so fast is that many companies already viewed operational work as lower-value toil, ripe for automation. Fred Hebert pointed out a revealing asymmetry in how AI tools are marketed: coding assistants are framed as augmenting the engineer (they’re “partners” and “teammates,” and the developer stays in control), while AI ops tools are framed as replacing the work entirely (“machines on-call for humans,” “stop firefighting, start innovating”). The framing reveals what the market thinks the work is worth: not much. If your company takes that view, that operational work is grunt work to be automated away, it’s going to underinvest in the human knowledge that effective incident response requires.

The operational work wasn’t just toil; it was keeping people’s heads in the game. Situational awareness gets built as a side effect of doing the work; you don’t realize its value until the work goes away and you discover the hard way that the awareness went with it.

The ironies of automation

In 1983, cognitive psychologist Lisanne Bainbridge published a paper called “Ironies of Automation” that described a paradox: the more you automate a process, the less aware the human operator is of the system’s current state, and the harder it becomes for them to handle the situations that the automation can’t. The most striking example is commercial aviation. Autopilot systems handle routine flight so well that pilots spend less time actively engaged with what the aircraft is doing. At the same time, pilots’ skills atrophy from disuse, since the autopilot is handling more and more of the work of flying. But when the autopilot fails or encounters something it can’t handle, the pilot needs to take over in exactly the kind of unusual situation that demands the most current awareness of the aircraft’s state and the most skill in responding to it.

There’s an old pilot joke that the scariest words you can hear in the cockpit are “Huh? What’s it doing now?” It’s funny because it captures exactly the gap Bainbridge described: the crew has lost track of what the automation is doing, at the moment when they need to understand it most.

This creates a double bind: the pilot is less aware of what’s happening right now, and over time, less practiced at handling it. The gap widens from both directions.

John Allspaw brought this concept to the DevOps and SRE community through his influential “A Mature Role for Automation” blog series, and it’s been shaping how we think about automation in software operations ever since. The principle isn’t anti-automation; it’s a caution about what automation displaces, and about what you need to do to compensate.

This is happening right now, fast

That pattern is now playing out with AI across operations, with one critical difference: the tempo. In aviation, the ironies of automation emerged over decades as new aircraft and autopilot systems were gradually introduced. In tech, traditional automation has been gradually eroding hands-on system knowledge for years. AI is compressing that same dynamic into months and weeks, because AI capabilities are advancing faster than any previous generation of automation, and because AI is automating categories of work that previous tooling couldn’t touch.

Consider an engineer whose team recently adopted AI tools for infrastructure management and code generation. Six months ago, they knew their deployment pipeline intimately because they built it, tuned it, and fixed it when it broke. They knew which services were fragile because they’d spent time troubleshooting them. They had a mental model of the system’s architecture because they’d worked with it directly. Now AI handles much of that work. The engineer is more productive. But when something goes wrong that the AI can’t resolve (when the situation becomes an incident), the engineer’s mental model is fuzzy and possibly six months stale. The discrepancy they would have noticed because they’d just been troubleshooting that service last week is now invisible to them.

And if you’re also using AI tools during incident response (as many companies are beginning to, for sitrep drafting, channel summarization, log analysis, and the like), the problem compounds: less system knowledge being brought into the incident, less situational awareness during the response itself.

Over time, the skills atrophy too. Engineers who rarely troubleshoot manually get worse at structured debugging. Engineers who rarely investigate anomalies lose the intuition for what’s worth pursuing. The immediate loss of awareness compounds into a longer-term erosion of skill, and as AI handles a wider range of situations, the situations that still require human judgment become harder and rarer. The humans facing those situations need to be more capable than before, not less.

What to do about it

None of this means companies should stop using AI ops tools. The productivity gains are real, and the direction is clear. The point is to deploy these tools with your eyes open about what the automation displaces, and to invest deliberately in maintaining the knowledge and skills that the automated work used to build.

The aviation industry recognized this problem decades ago. Their solution: frequent mandatory recurrent training, specifically designed to keep pilots practiced on the skills that routine automated flight no longer exercises. The tech equivalent is exercises, game days, and deliberate hands-on work with the systems your team is responsible for. Companies like Uptime Labs are starting to build tools for exactly this kind of recurrent training.

Keep engineers connected to the systems they’re responsible for. If AI handles most of the day-to-day operational work, create deliberate opportunities for engineers to work with the systems directly: manual deployments during low-risk windows, hands-on troubleshooting during exercises, periodic deep-dives into the infrastructure that go beyond what the AI dashboards show.

Treat AI tools as useful but not required. Build your processes so that an AI tool outage is an inconvenience, not a crisis. The engineers who can still function without the tools are the ones you’ll need when the tools aren’t available (and they won’t be, eventually; tools fail, sometimes during the incidents where you need them most).

Run exercises and game days that test system understanding, not just process compliance. A tabletop exercise where the scenario is “your AI ops tools are down and you need to investigate a production issue manually” will tell you a lot about how much system knowledge your team has actually retained.

The bottom line

The ironies of automation aren’t an argument against automation. Bainbridge wasn’t arguing against autopilots, and this isn’t an argument against AI ops tools. The argument is that automation changes the human’s relationship to the work in ways that are easy to miss. The operational work wasn’t just toil; it was building the understanding that people need when things go wrong. When AI takes over that work, the understanding erodes, quietly and steadily, until the next incident reveals how much has been lost. Companies that deploy AI ops tools without accounting for this will discover the gap at the worst possible moment: during the incident that the AI can’t handle, when the engineers discover they’re no longer ready to handle it either.


I’m writing a book on incident management for DevOps and SRE that covers this and much more. Sign up at im4ds.com to be notified when it’s available.

If your company needs help with incident management right now, my consulting practice is greatcircle.com/im.

The incident metrics mirage

Most companies that invest in improving their incident management see something counterintuitive in the first few months: their incident count goes up. And someone in the leadership chain, looking at the dashboard, asks with concern: “Why are things getting worse?”

The good news is, they’re probably not. What you’re really seeing is evidence that your incident culture is getting better and stronger.

Incident count doesn’t measure system health; it measures how willing people are to declare an incident. When you invest in your incident management program by giving people better training, tools, and processes for handling incidents, they put those things to work, and more situations get treated as incidents. Problems that used to be handled informally (a “spicy bug” that someone handled without declaring an incident, a degradation that the on-call engineer white-knuckled through without telling anyone) now enter your incident process. The company is getting visibility into problems that used to go unnoticed, and using incident management practices and tools to address problems it would have struggled with before.

And it’s a self-reinforcing cycle; the more people use these skills and tools, the more comfortable they get with them, and the more they’re inclined to use them. That’s a good thing, as the way to build the skills and confidence to handle big incidents is by handling lots of little ones; the small incidents are an invaluable training ground.

In most companies, you’re dealing with increases in multiple dimensions simultaneously: number of users, number of products, number of features in those products, usage of those features, number of engineers, level of training and experience of those engineers, and many more. With all those factors increasing, why would you expect incident count to decrease? In a very real sense, rising incident counts can be a sign of success, not failure.

The question to ask isn’t “why are we having more incidents?”, it’s “why aren’t we?” And from there, other good questions follow: are we handling our incidents well, do we have the tools and training that we need, are we learning everything we can from every incident, are we preventing future incidents, are we better prepared to handle those we can’t prevent?

But what happens if leadership decides that rising incident count is a problem and sets a target to bring it down? People get the message: fewer incidents is better. So marginal incidents stop getting declared. The “spicy bugs” go back to being handled quietly. The degradations get white-knuckled through again. The number on the dashboard goes down, but the company has lost both the visibility its incident process was providing and the benefits of handling those situations with proper coordination, communication, and prioritization. The problems didn’t go away; you just stopped applying your best tools to them.

The Goodhart’s Law problem

This pattern has a name: Goodhart’s Law. When a metric becomes a target, it ceases to be a good metric. And it’s not just an incident count problem; it recurs across every incident metric companies reach for.

The logic is straightforward. You focus on a particular metric because you think it captures something you care about. You set a target for the metric because you want to improve. And then smart, well-intentioned people find ways to hit the target. Some of those ways involve actually improving the thing you care about. But some involve optimizing the number without improving the underlying reality, and over time, the second category tends to dominate.

Make incident count a target for reduction, and marginal incidents stop getting declared. Make MTTR a target, and people close incidents prematurely. Make action item completion rate a target, and people write easy action items instead of hard ones. In every case, the metric looks better while the thing you actually care about (learning, reliability, preparedness) stays the same or gets worse.

This is human nature. Smart people optimize for what gets measured; that’s how incentives work. You’re not going to prevent it by writing sternly worded memos about gaming the system. You can only manage it by choosing metrics carefully and then paying attention to the behaviors driven by your focus on those particular metrics. Adjust when those behaviors aren’t what you intended, and be willing to retire the metric when it’s doing more harm than good.

MTTR: the metric everybody loves and nobody should trust

Mean Time to Recovery is the most misleading incident metric in the industry. Leadership loves it because it’s a single number that appears to capture “how fast we fix things.” But it’s deeply flawed, both mathematically and in the incentives it creates.

Incident durations follow a power-law distribution: most incidents resolve quickly, while a small number take much longer. When you average power-law data, you get a number that describes nobody’s actual experience. Let’s imagine that last month you had ten incidents, nine of which resolved in about ten minutes each, and one that took six hours. That gives you an MTTR of 45 minutes, but that’s nowhere close to what any of those incidents actually took; it’s way off of both ten minutes and six hours.

Google SRE ล tฤ›pรกn Davidoviฤ, in Incident Metrics in SRE: Critically Evaluating MTTR and Friends (O’Reilly, 2023), used Monte Carlo simulations to demonstrate that even with a substantial dataset, MTTR can’t reliably tell you whether your incident response is actually improving. The math doesn’t just give you a misleading number; it can’t even detect real improvement when it’s happening.

There’s also a flattening problem: MTTR treats all incidents as interchangeable, as if the only thing that matters about an incident is how long it took. A six-hour incident where page load times were degraded but the service was still usable somehow scores worse than a one-hour total outage.

The incentive problems are even worse than the mathematical ones. MTTR incentivizes speed over understanding. Thorough incident response sometimes means deliberately slowing down: carefully analyzing symptoms, verifying that a fix actually works, understanding the contributing factors well enough to prevent recurrence. MTTR punishes all of that. A team that’s genuinely improving (catching issues earlier, preventing cascades, tackling more complex problems) can see its MTTR stay flat or even go up. The metric undermines morale and leadership confidence even as the team does better work.

What fire departments get right about metrics

Here’s a lesson that most software companies could learn from fire departments.

Well-managed fire departments decompose their response timeline into segments and set targets only on the segments they can actually control, and expect to be fairly consistent across their incidents. Dispatch time (how long from answering the 911 call to notifying the fire crew) gets a target. Turnout time (how long from notification to crews leaving the station) gets a target. Drive time (from the crews leaving the station to arrival at the incident scene) gets a target. These are process steps that are largely consistent from one call to the next, and if they’re too slow, you can do something about it: hire more dispatchers, change how crews stage at the firehouse, build more stations.

What fire departments don’t set targets for is how long it takes to put out the fire. That depends on the fire. A dumpster fire and a fully involved warehouse fire are different problems with different durations, and no fixed target could meaningfully apply to both.

How long it takes to set up an incident channel, how long it takes responders to acknowledge pages, how long it takes the responders to join the channel and get started: these are your equivalent to the fire department’s dispatch chain. They’re largely consistent from one incident to the next, and you can set targets for them. If you’re not meeting the targets, there are obvious adjustments you can make: auto-create incident channels, set and enforce clearer on-call expectations, and so forth.

On the other hand, the time to find the contributing factors, the time to implement a durable fix, and the total time for the incident: those depend on the details of the particular incident. A misconfigured feature flag and a cascading database failure are different problems. Track trends, investigate outliers, learn from reviews. But don’t set targets. Targets on metrics you can’t control produce gaming, demoralization, or both.

Start with questions, not metrics

The most useful thing is to stop asking “what should we measure?” and start asking “what questions are we trying to answer?”

Are incidents being handled well? Are we learning from them? Is the incident management program serving the business? Is our on-call workload sustainable? Each of these questions leads you to seek different evidence, some quantitative, some qualitative, and the answers are more useful than any single number on a dashboard.

These aren’t easy questions to answer, but they’re better than an easy number that misleads you (like MTTR).

Your dashboards should make you curious, not confident. When you see a trend, the right response isn’t “we know what’s happening.” It’s “we should dig in and find out why.”

And if your incident count went up this quarter? Before you panic, ask why. You might find that your increased focus on incident management is doing exactly what it’s supposed to do.


This is one of the topics I cover in depth in my upcoming book, Incident Management for DevOps and SRE. If you’d like to hear when it’s available, you can sign up at im4ds.com.

If your company needs help with incident management right now, my consulting practice is GreatCircle.com/im.

Most incident “freelancing” is really an infrastructure problem

Your team burned 40 minutes in an incident chasing a ghost in the metrics. It turned out someone unaware of the incident had picked a bad time for a routine restart.

This is sometimes called “freelancing,” or “going rogue.” It’s working on (or near) the incident without being part of the organized response, and it’s one of the most common complaints I hear from engineering leaders when they talk about incident management: “Our people go off and do their own thing instead of coordinating.”

And they’re right to be concerned. Uncoordinated work during an incident is genuinely costly. Incident “freelancers” muddy the trail. Their queries bog down log systems, making other responders’ searches slower. Their investigations generate artifacts that get mistaken for symptoms of the actual problem. They make changes that mask the issue or introduce new ones. In the scenario above, the incident responders lost 40 minutes because someone’s routine restart looked, in the dashboards, like a clue.

The instinct is to treat this as a behavior problem: tell people not to freelance, put it in the incident guidelines, remind everyone in training. And if they keep doing it, escalate.

Unfortunately, in my experience, most of what gets called “freelancing” isn’t actually freelancing.

What real freelancing looks like

Real freelancing is a deliberate choice: someone knows an incident is underway, knows there’s an organized response, but decides to work the problem independently anyway. Maybe they think they’ll be faster on their own. Maybe they don’t trust the incident commander to use them effectively. Maybe they’ve had frustrating experiences in past incidents and decided it wasn’t worth the trouble.

This is a real phenomenon, and it’s worth taking seriously. When someone with relevant expertise actively avoids the coordinated response, that’s a signal about your incident management culture: something about the experience of participating is broken enough that a skilled person would rather work alone.

Fortunately, this kind of deliberate freelancing is rare. Most companies that think they have a freelancing problem actually have something quite different.

The three gaps

Go back to the ghost in the metrics caused by the routine but uncoordinated restart. The person who did the restart wasn’t working the incident independently; they weren’t working the incident at all. But why were they doing a routine restart in the middle of a Sev-1?

Maybe they had no idea an incident had been declared. Maybe they knew something was going on but didn’t think it involved their systems. Or maybe they suspected it might be a bad time, but had no way to check. Same outcome, three very different backstories, and the fix is different for each one.

When you look at the incidents where uncoordinated work caused problems, most of them trace back to one of these three structural gaps.

The visibility gap. The person didn’t know an incident had been declared. Maybe the declaration went to a channel they aren’t in, or one they haven’t caught up on yet. Maybe the alerting didn’t reach their team. Maybe they were heads-down in focused work and missed it entirely. They weren’t choosing to work outside the response; rather, they didn’t know there was a response to join.

This is the most common gap, and it’s the one that produces the most collateral damage, because the person has no reason to think their normal work might interfere with anything. They restart a service, run a migration, deploy a config change, all routine, all uncoordinated with the incident responders, and all potentially confusing to responders trying to interpret what they’re seeing in the dashboards.

The identity gap. The person is aware something is going on but doesn’t see themselves as relevant. “That’s a payments incident; I’m on the search team.” They carry on with their normal work, not realizing that both teams depend on the same cache cluster and that the real problem is there. The incident responders don’t know to warn them, because they don’t see the shared dependency either.

This gap is subtler than the visibility gap. The information about the incident reached the person; the connection to their own work didn’t.

The mechanism gap. The person knows about the incident and suspects their work might be relevant, but there’s no clear way to check. There’s no place to ask “is now a bad time for routine changes?”, no lightweight way to raise a hand and coordinate. So they make a judgment call, usually in the direction of “it’s probably fine,” and carry on.

This is the gap that frustrates well-intentioned people the most. They would have coordinated if there had been a well-understood way to do so. But the response didn’t have one, so they did the best they could with the information they had.

The plumbing fix

Unlike “real” (i.e., intentional) freelancing, all three of these gaps are plumbing problems, not people problems. They’re about whether your incident response infrastructure makes it easy for people across the company to know an incident is happening, to understand whether their work might be affected, and to coordinate without joining the full response.

What would that look like? In the everyday world, the flashing lights at an emergency scene are a broadcast signal to everyone in the vicinity. They tell passing drivers, pedestrians, and nearby work crews: something is happening here, adjust your behavior. Nobody expects individual firefighters to personally flag down every car that might drive through the scene. The lights do that job passively, at scale, without coordination.

The flashing lights can also serve responders. At an incident scene (in the US, at least), a green flashing light marks the command post, where the incident commander can be found, and arriving responders know to go there to check in for an assignment.

Most companies don’t have the equivalent of flashing lights for their incidents. The incident declaration establishes the incident channel, and the paged responders join it, while everyone else in the company carries on unaware.

The fixes are concrete and mostly unglamorous, and include:

Broad incident visibility. When an incident is declared, the notification should reach beyond the directly-paged responders. A company-wide incidents channel, automated cross-posts to team channels for affected services, a banner in internal tools: whatever fits your company’s communication patterns. The goal is that anyone doing work that might intersect with the incident has a reasonable chance of knowing about it. For significant incidents, that notification can include a simple “hold non-urgent changes, or check with the incident channel” signal.

Dependency-aware notifications. When an incident is declared for one service, teams that own connected services often don’t realize the incident might involve them. If your company maintains a service dependency map (even a rough one), use it: automatically notify teams whose services are upstream or downstream of the affected system. “Heads up: there’s an active incident involving the payments service, which depends on your cache layer” turns “not my problem” into “maybe I should hold off on that restart.”

A lightweight coordination path. Not everyone who might be affected needs to join the incident response. But they need a way to check in: “I was about to restart the cache fleet; is that going to cause problems for you?” A cultural norm that it’s OK (even expected) to ask in the incident channel turns invisible collisions into two-minute conversations.

These aren’t expensive changes. They’re the kind of infrastructure that, once built, quietly prevents dozens of 40-minute detours a year. They work because they address the actual problem: most people who end up doing uncoordinated work during incidents aren’t choosing to go rogue. They just didn’t have the information or the path to coordinate.

Real freelancing, the deliberate kind, still deserves attention. But if you’re seeing a pattern of uncoordinated work during your incidents, start with the plumbing before you start with the lectures.


I’m writing a book about incident management for software engineering companies. If you’d like to hear about it when it’s available, sign up at im4ds.com. And if your company needs help with incident management right now, my consulting practice is GreatCircle.com/im.

The Post-Incident Review Meeting: three meetings in a trench coat

Most companies hold a single post-incident review (PIR) meeting for each incident. They schedule an hour, invite the responders and a handful of observers, walk through the timeline, discuss what went wrong, generate a list of action items, and move on. It feels productive. The calendar invite says “PIR Meeting,” and the meeting does PIR-meeting things.

The purpose of a post-incident review is to learn. Not to assign blame. Not to generate action items. Not to produce a document for the compliance folder. Documentation and action items are side effects of the review process, but they aren’t the point. The point is learning, both individual and organizational, so that you have a better understanding of how your systems actually work and you’re better prepared for the next incident. Because there will be a next incident.

But what most companies call “the PIR meeting” is really three distinct functions crammed into a single calendar invite (or, as I’m fond of describing it, “three meetings in a trench coat”):

A working meeting where the people who responded to the incident sit down together and reconcile their understanding of what happened. They fill gaps in the timeline, surface things they knew but didn’t write down, and pressure-test the contributing factors.

An action items meeting where problems that the incident surfaced are named. The goal should be to identify what needs attention, not to propose solutions; the people best positioned to design fixes may not even be in the room.

A presentation where the findings and lessons are shared with a broader audience beyond the people who lived the incident. This is the meeting’s contribution to organizational learning: spreading what was learned to people who weren’t in the room.

Each of these three functions has a different optimal participant list, a different facilitator posture, a different conversational mode, and a different relationship to time pressure. The working meeting is a small group, collaborative and sometimes messy, exploring what happened without a fixed agenda. The action items meeting shifts into problem-identification mode: what did this incident reveal that needs attention? The presentation is structured and scripted, aimed at an audience that wasn’t in the room for the incident.

If you treated them as three separate meetings, you’d probably invite three different (though overlapping) sets of people to them. Which means that in a single combined meeting, either you haven’t invited everyone who should be there for each function, or some of the people you invited are sitting through parts of the meeting that they don’t need to.

Also, the three functions aren’t equally important: the working discussion is the foundation that the other two depend on. When they’re collapsed into a single session, the results are predictable.

What happens when the three functions compete

When these three functions share a single meeting, the action items tend to dominate, because “what are we going to do about this?” feels like a more urgent conversation than “what can we learn from this?” Especially under time pressure, the room gravitates toward the concrete and seemingly actionable, at the expense of the exploratory and uncertain. Someone says “we should add monitoring for this,” and the conversation shifts from understanding what happened to debating what to build. Once that shift happens, it’s hard to get back.

Meanwhile, the broader audience sits passively through a working discussion that wasn’t designed for them. The observers are theoretically “learning,” but when the conversation is a detailed working discussion among the responders, observers tend to drift to Slack and email, half paying attention at best. The presentation function doesn’t just get less time in a combined meeting; it gets less attention.

And the tyranny of the one-hour calendar block hangs over everything (especially if your hour only has 53 minutes). The working discussion goes where the work takes it. You can’t predict how long it needs based on the severity or complexity of the incident. Sometimes there’s a lot to learn from a small incident; sometimes there’s surprisingly little to discuss about a big one. A one-hour combined meeting trying to do the work of three distinct functions will likely shortchange all of them.

A diagnostic, not a prescription

The three-meeting frame isn’t a prescription to hold three meetings for every incident. Even at companies with the most mature incident practices, most incidents get a single meeting. That’s fine.

The frame is a diagnostic tool. If your review meetings feel rushed, performative, or dominated by action items, the problem might be that you’re asking one meeting to do the work of three. Recognizing the three functions helps you protect the one that the other two depend on (the working discussion) when they share a calendar invite, and invest in separate meetings for the incidents that warrant it.

Protecting the learning conversation

Understanding without follow-through is just conversation. But follow-through without understanding is just busywork. The action items that come out of a review are only as good as the understanding that produced them.

When you do hold a combined meeting, the simplest tool for protecting the learning conversation is to explicitly defer discussion of action items. “We’re going to defer discussing action items until the last fifteen minutes. If something comes up that feels like an action item, note it and we’ll come back to it.” Then enforce the boundary. During the working session, when someone says “we should add monitoring for this,” the facilitator says “noted; write that down so we can come back to it.” The first few times feel awkward. It gets easier, and the room gets better results.

In my experience, which is shared by other leading practitioners in the LFI (learning from incidents) community, teams generate fewer and better action items when they let the understanding generated in the working session settle and marinate a bit before they start digging into “what needs to change?” When people have time to sit with the understanding before jumping to “what are we going to do about this?”, they move past the reactive fixes and toward improvements that address broader patterns. If possible, you should defer the action items discussion until 24 hours after the working meeting; that’s not always practical, but the separation produces higher-quality outcomes.

The pattern is older than software

Separating these functions has parallels in other fields. The NTSB (the U.S. National Transportation Safety Board) separates its investigation from its public hearings from its final recommendations. Hospitals hold weekly morbidity and mortality conferences where cases are presented to the broader department, informed by detailed case review that happens separately. In both fields, the investigation, the discussion, and the recommendations are distinct phases. The principle applies whether you’re investigating a plane crash, a surgical complication, or a database outage.

None of this is exotic or expensive. It’s a matter of recognizing that a single meeting is trying to do three different jobs, naming those jobs, and deciding which one matters most when they compete.

For most incidents, this means giving the working discussion room to breathe, and keeping action items from taking over before understanding has had a chance to develop.


I’m writing a book on Incident Management for DevOps and SRE. Sign up to be notified when it’s available.

If your company needs help building or improving its incident management capabilities, my consulting practice is Great Circle.

The Problem with AI-Generated Post-Incident Reviews

Modern AI tools can produce a competent-looking post-incident review document from a Slack channel transcript and a few prompts. The output will be pleasingly formatted, with a timeline, a list of contributing factors, and a set of action items. It will read coherently, and it will arrive faster than a human-written review would have, with less engineer time spent producing it. For a manager looking at the post-incident review process and seeing engineers grumble about the time it takes, this is tempting.

The catch is that the document was never the point of the review. The real learning comes from analyzing the incident while writing the document, not reading it; the document at the end is the residue of the learning. It’s like studying; you learn a lot more from working the problem sets than you do from just reading a classmate’s summary.

Three layers of learning

The learning happens at three layers: readers of the published document, the writers individually, and the writers as a group.

Readership is the most visible layer, and radiates outward when the document is published. Colleagues throughout the company read the review, see how it says the system behaved and how the team supposedly handled the incident, and perhaps update their own understanding of what’s possible. Most of them weren’t in the incident, so for them, the document is the incident. Their learning is downstream of the writers’ analysis, and weak analysis produces shallow lessons; readers get less than they could have, and some of what they get may be flat-out wrong.

Each writer ends up with a more thorough understanding of the incident than they started with. They start to write “the deploy caused the outage” and realize, as they trace the sequence, that the deploy only surfaced a problem that was already lying in wait. They find themselves describing the dashboard as “down” and stop, because the dashboard wasn’t actually down; it was up, but showing data from the wrong cluster, which is why nothing made sense to the responders for the first eighteen minutes. They write “the team decided to…” and stop, because the team didn’t decide; one person made a call and the others went along, and the gap between those two things turns out to matter.

The writers also learn from each other. Responders fill in their slices of the timeline; the architect annotates the contributing factors; customer success writes the impact section. As they go, someone catches a gap, someone corrects a misremembered moment, a disagreement surfaces in the comments and leads to an enlightening discussion. By the time the writing is done, the group has a grasp of what happened that no individual writer had alone. They built it from reconciling what each of them separately knew.

What AI does to each layer

When AI writes the document, each of these layers fares differently, and none of them fares well.

Readers are still readers. They open the document, take it in, and maybe update their understanding based on what they read. But what they’re absorbing now is the AI’s synthesis, with no human pressure-testing behind it. It may or may not be right, or get to the deeper issues that a group of writers might have uncovered. The document looks like a review, and readers absorb it like a real review. Whether they’re learning anything true or useful is now a function of how well the AI happened to do, with no way to tell from the outside.

There are no writers when the AI does the writing, so there’s no individual learning from the writing process. No one stops mid-sentence to discover that the deploy only surfaced a problem already lying in wait; no one finds that the dashboard was up but showing wrong-cluster data; no one writes “the team decided to” and stops to reconsider. The kind of learning that comes from working through the evidence sentence by sentence doesn’t happen when no one is doing it.

And without writers, you clearly can’t have “writers learning from each other.” The AI conjures up a plausible-sounding narrative from whatever it was given. No one catches a gap; no one corrects a misremembered moment; no disagreement surfaces in the comments to lead to an enlightening discussion. Any tension between what different people actually thought is gone before anyone could surface it.

So you end up with a polished artifact. Your contributing factors are the AI’s guess at what’s plausible, not your team’s hard-won understanding. Your timeline is a transcript reorganization, not a reconstruction. Your “lessons learned” come from the AI pattern-matching against incidents in its training data, not against the ones your company actually had. The document looks like a review, but it’s a fantasy.

You can automate production of the review document. You can’t automate the understanding that the process was supposed to produce. Automating the writing away automates the learning away.

Where AI legitimately helps

This isn’t a case for keeping AI out of the post-incident review process. Rather, it’s a case for being clear about where AI helps the engineers do the work and where it does the work for them. Mechanical support is genuinely useful. Substitution for human thinking is not.

Specific places AI earns its keep:

  • Collating raw material. Pulling content from multiple Slack channels and threads into a unified view, transcribing voice channels (either real-time or recorded), gathering the screenshots and graphs that got posted in-channel during the incident. The grunt work that gives writers a clean starting point.
  • Format and copy editing on a draft a human has written. Tightening prose, suggesting clearer phrasings, catching the inconsistencies that survive a human edit pass. AI is good at this, and using it here saves time without taking anything away from the writer’s engagement with the material.
  • Surfacing gaps. “Your timeline jumps from 14:46 to 15:12 with no entries. Was anything happening then?” That’s a useful prompt for human writers, not a substitute for them.
  • Cross-referencing past incidents. “Three other incidents in the last six months touched the same service; here are the links.” Pattern-matching across a library of past reviews is exactly the kind of mechanical work AI is well-suited to. (This one requires that your library of past reviews actually exists and is structured in a way the tool can search, which is a separate problem worth solving on its own merits.)

AI handles mechanical work that supports the writers. The writers do the thinking. The moment you let the tool do the thinking (generating the contributing factors, drafting the lessons learned, writing the narrative itself), you’ve automated away most of the learning, and arguably the most valuable parts.

The right question for the manager

If you’re an engineering leader evaluating an AI tool that promises to help with post-incident reviews, the question isn’t “does this tool produce a good-looking document faster?” All of them do that, but appearance alone shouldn’t be your criterion. A much better question is “does this tool support my engineers in doing the writing, or does it replace them in doing it?”

Tools in the first category save real time on the parts of the work that don’t produce learning. Tools in the second category save your engineers from the activity that the post-incident review exists for. They give you a polished artifact and an empty experience. The same incidents keep happening, but hey, now you have a faster pipeline for marking off the “Write Post-Incident Review” checkbox.

Here’s a clarifying way to think about it: you could throw the post-incident review document away after writing it and still get the vast majority of the value out of the process. The document is like the scribblings on a whiteboard after a productive working session: interesting, yes, and maybe worth snapping a photo of, but the real value is what leaves the room in the heads of the people who were there. You don’t actually want to throw it away (the document does real work, both immediately and over time, as part of the library of past reviews), but knowing that you could is the right reference point for thinking about which tools genuinely help the process and which ones quietly hollow it out.

AI can produce a good-looking incident review document, but only your engineers can produce the understanding behind a truly good one. Adopt tools that support that distinction; the ones that don’t will leave you with a stack of polished artifacts, but without much actual learning.


Post-incident reviews are one of several topics covered in my forthcoming book, “Incident Management for DevOps and SRE.” Learn more and sign up for updates at im4ds.com.

Need help preventing, preparing for, responding to, and learning from incidents? That’s the focus of my consulting practice at Great Circle.

Self-blame isn’t blameless

Most of what’s been written about blameless post-incident reviews is about managers not blaming engineers, and engineers not blaming each other, because blame shuts down learning. What many miss is that engineers still blame themselves, and the damage is the same.

The most valuable output of a post-incident review isn’t the document, and it isn’t the list of action items. It’s learning: individual learning for the people involved, and organizational learning that outlasts anyone’s tenure. Everything else in the review process serves that goal.

The scene

Somewhere in the middle of a post-incident review, one of the engineers involved in the incident speaks up:

“Yeah, I should have caught that. My bad. I’ll be more careful next time.”

On the surface, this feels like maturity. The engineer isn’t defensive; they’re owning what happened. The team nods, accepts the explanation, and the review moves on.

What actually just happened is that the exploration stopped. “Sam made a mistake” became the story, and all the systemic factors that contributed to the incident went unexamined. The dashboard that was slow to load. The deployment pipeline that didn’t have a canary phase. The runbook that was last updated eighteen months ago. The schedule pressure that made skipping the check seem reasonable in the moment. None of that got talked about, because “I’ll be more careful next time” was accepted as the explanation, when it doesn’t really explain anything.

Self-blame is still blame

Blame shuts down learning. That’s true whether the blame comes from a manager, from a peer, or from self-blame. A satisfying-sounding explanation arrives early in the conversation, everyone accepts it, and the exploration that would have surfaced the actual contributing factors never happens. The document captures a neat narrative. The action items address the narrow issue. The next incident reveals that the narrow issue was a symptom of something deeper that nobody investigated.

“I should have caught that” is a thought-terminating clichรฉ that sounds like insight, but isn’t. It works just as well whether someone else says it or the person says it about themselves.

The heroism problem after the incident

If your company has spent any time thinking about incident response, you’ve probably internalized a version of this principle: individual heroics are a problem during incidents. The lone-wolf engineer who goes heads-down and fixes the outage by personal brilliance and sheer force of will is not the model you want. The incident runs longer than it needed to, nobody else learns anything, and the company ends up dependent on a handful of heroes who burn out or leave. Organizational capability beats individual heroics.

Self-blame is the same pattern, just moved later in time.

The engineer who says “my bad, I’ll be more careful next time” is volunteering to carry the weight of the incident themselves. They’re taking the hit so the team can move on. It feels selfless, even noble. It also deprives the team of the conversation it actually needs to have about the contributing factors behind the incident. The hero absorbs the cost, and the company never learns what it needed to learn.

The response to heroism during incidents is to build structures that don’t require it: defined roles, clear handoffs, explicit delegation, training that spreads capability across the team. The response to heroism after incidents should be the same: don’t let one person carry (and bury) what needs to be shared.

What to do when you see it

Self-blame shows up in a handful of recognizable forms. “My bad, I’ll be more careful.” “I should have known better.” “I take full responsibility for this.” “That was on me.” Whichever variant surfaces, the review needs someone else to gently redirect. Not to argue with the person, and not to wave them off. The work is to pull the conversation from the person’s character back to the situation they were in.

A question that works: “OK, but let’s understand the circumstances. What were you seeing at the time? What information did you have? What made the action seem like the right thing to do?”

This question does several things at once. It signals that the team isn’t satisfied with the self-blame as the explanation. It treats the engineer’s actions as reasonable given what they knew, which is almost always the case. And it opens up the conversation about the circumstances the engineer was working in: the missing information, the ambiguous signals, the process that everyone knew was broken but nobody had fixed. That’s where the contributing factors live. That’s where the learning lives.

Facilitators of post-incident reviews should anticipate self-blame, and be prepared to address it.

The person in the middle

There’s another aspect of this that seldom gets the attention it deserves: the toll on the person absorbing the blame.

The engineer most involved in the incident often arrives at the review already carrying guilt and anxiety; the “my bad” moment is frequently the release of pressure they’ve been carrying for days. Treating it with care, neither accepting the self-blame as the answer nor challenging the person, is part of the job.

The burden doesn’t ease when the review ends. The engineer absorbing the blame carries it for a long time afterward. In healthcare, this is called “second-victim syndrome”: dealing with a patient’s trauma sometimes causes adverse emotional effects for the healthcare provider, as well, making them the “second victim.”

The stakes in software are different, but the experience isn’t as different as it might sound. Engineers lose sleep over incidents. They dread going to work. Some of them leave their jobs. “I should have caught that” sounds like a quick self-assessment; in practice, it’s often a thought the person has been rehearsing for days before the meeting.

A well-run review that explores the systemic factors behind the incident is one of the most effective interventions for this. It reframes the experience for the person involved. They’re no longer “the person who caused the outage.” They’re a participant in a systemic event the company is working to understand and prevent. That reframing matters enormously, and it can’t happen if the review stops at accepting the self-blame as the answer.

Variations to watch for

Individual self-blame has relatives worth watching for. Group self-blame sounds like “we should have caught that”; it diffuses the ownership across the team but stops the conversation the same way. Passive voice (“the change was deployed without adequate testing”) strips the actor from the sentence without removing the judgment. “Human error” without naming the human is blame at one remove, assigned to a generic stand-in. All of these give the review a stopping point that feels satisfying but isn’t actually useful. The redirect is the same in each case: pull the conversation from implicit character judgments back to the situation people were in.

The principle

The next stage of blameless maturity isn’t arriving at a review where nobody points fingers. It’s arriving at a review where the conversation keeps going past “my bad, I’ll be more careful next time” and into the system that made the moment possible.

Fred Hebert has written about a related anti-pattern he calls “superficial blamelessness“: reviews that successfully avoid retribution but still land on individualistic fixes (more training, pay more attention, add supervision) rather than changes to the system. Self-blame is a particularly sneaky version of that pattern. The engineer volunteers the individualistic remediation on themselves, which makes it feel like accountability instead of the shallow fix it is.

Accountability means understanding how your company and its systems actually operate, and making durable changes based on that understanding. One person promising to try harder doesn’t get you there.


I’m writing a book, Incident Management for DevOps and SRE, aimed at helping companies build incident management capability that doesn’t depend on heroics. Sign up for updates at im4ds.com.

If your company needs help with incident management right now, my consulting practice is at greatcircle.com/im.

Why you keep losing your incident commander to debugging

Ninety minutes into an outage, the person who’s supposed to be running the incident has lost the plot. They’re hunched over a terminal, or deep in a dashboard, or engrossed in a Slack thread debugging the problem alongside the responders. Nobody’s sending status updates. Nobody’s fielding questions from the support team. Nobody’s thinking about whether the response needs more people, or different people, or a completely different approach. The incident commander (IC) has disappeared into the technical work.

When this happens (and it happens constantly), companies tend to treat it as a discipline problem. “The IC got sucked in again.” “We need ICs who can stay above the fray.” As if the solution were stronger willpower.

This is a cognitive problem. Technical troubleshooting and incident coordination require fundamentally incompatible kinds of attention, and doing both at once means neither gets the attention that it needs.

Two rhythms that don’t mix

Technical troubleshooting has a rhythm. It’s heads-down, focused, and sustained, even when you’re working the problem with other responders. You’re following a thread: correlating timestamps, forming a hypothesis, testing it, adjusting, testing again. The work rewards concentration. Interruptions are expensive; every time you break your focus, you lose the mental model you’ve been building and have to reconstruct it. The best troubleshooting happens when you can tune everything else out and just chase the problem.

Incident coordination, the incident commander’s actual job, has the opposite rhythm. It’s heads-up, scanning, bouncing from thing to thing. “Has the networking team checked in?” “Where’s that status update?” “Do we need to loop in the database on-call?” “The VP of engineering just asked for an ETA.” The work requires constant context-switching. Staying focused on any one thread for too long means everything else drifts. If the incident commander spends ten minutes deep in a technical discussion, they’ve missed three stakeholder questions, a newly joined responder has no idea what to work on, and nobody outside the response has heard anything since the incident started.

These two rhythms don’t just coexist poorly, they actively fight each other. The focus required for effective debugging is exactly what makes for ineffective coordination. The constant interruptions that are part and parcel of coordination are exactly what make for ineffective debugging.

The gravitational pull of debugging

When one person tries to do both, the technical work almost always takes over. This isn’t surprising. The technical work is tangible, intellectually engaging, and feels more immediately productive. You’re making progress, finding clues, narrowing the problem. The coordination work is less satisfying in the moment. Writing a status update doesn’t feel like fighting the fire, it feels like paperwork.

So the IC drifts. They open a dashboard “just to check something.” They start a query “just to confirm a hunch.” Twenty minutes later, they’re deep in the investigation and the coordination work has stopped entirely. Nobody told them to stop being the IC, and they didn’t make a conscious decision to stop. They just drifted away from it, because the technical problem was right there and it was interesting and they could help.

The people around them usually won’t say anything, either. The responders are glad to have another strong technical mind on the problem. The stakeholders waiting for updates assume someone is handling it. By the time anyone notices that coordination has stopped, the damage is done: stakeholders are confused, new responders have self-dispatched to random tasks, and nobody has a clear picture of the overall situation.

What to do about it

The fix is structural: don’t expect one person to both lead the response and do technical work within the response.

That expectation often comes from the IC themselves. A strong engineer who’s “just coordinating” can feel like they’re not pulling their weight, especially when they can see exactly what needs to be tried next. If the IC genuinely has unique knowledge that the response needs (they built the failing system, they’ve seen this failure mode before), the right move is for them to hand the IC role to someone else and join the response as a responder. Trying to do both isn’t a good solution.

Think of an orchestra conductor. The conductor doesn’t play an instrument; the orchestra as a whole is their instrument. The moment the conductor picks up a violin, nobody’s conducting. The same thing happens when an IC opens a terminal.

The incident commander doesn’t debug the outage itself; they debug the incident response.

The IC stays in the coordination rhythm: tracking the response, communicating outward, and keeping the big picture in mind so they can make what are often called “sacrifice decisions.” Should we sacrifice the last hour of customer data to roll back to a known good state? Should we keep the storefront down and fix the problem properly, or bring it back up in a degraded state and risk a second outage? Should we notify customers now with incomplete information, or wait until we know more? The technical team can lay out the options, but these tradeoffs cut across teams and affect the business in ways that someone heads-down in a terminal can’t see.

On a small incident (i.e., one involving the IC and just one or two responders), you don’t need any formal role designations among the responders; you just need a clear split between coordinating (by the IC) and troubleshooting (by the responders). On bigger or more complex incidents, where several responders are working the problem, it often makes sense to designate one of them as the tech lead for the response.

The IC and tech lead roles face in opposite directions. The IC faces outward, toward the rest of the company: stakeholders, executives, support teams, other engineering groups that might be affected. The tech lead faces inward, toward the problem: directing the investigation, synthesizing what the responders are finding, making the tactical calls about what to try next. Each one watches a different part of the horizon, and together they cover the full picture.

The IC and tech lead stay in close touch with each other. The tech lead gives the IC a clear, concise summary of where the investigation stands and what the team needs. The IC keeps stakeholders informed, and makes sure the technical team has what it needs to keep moving: resources (including additional responders, if needed), priorities, and decisions on the tradeoffs that aren’t visible from inside the investigation. Together, they let the investigation go deep without the response going dark. Neither one has to switch cognitive modes. Each stays in the rhythm that makes them effective.

When the IC stays in the coordination rhythm, the scene looks different. Stakeholders are getting updates. New responders know what to work on. The responders are heads-down, uninterrupted, chasing the problem. And nobody had to be told to “just try harder.”


I’m writing a book on incident management for engineering teams. If this resonates, visit im4ds.com to follow along.

If your company is working through challenges like this one, I do consulting and training on incident management for engineering companies.

Why NASA flight director Gene Kranz is the gold standard for incident commanders

Who was the greatest incident commander of all time? My money is on Gene Kranz.

If you’ve seen the 1995 movie Apollo 13 (and if you haven’t, you should; it’s a gripping tale even though you know the outcome, and along the way it’s a master class in incident command), you’ve seen Ed Harris portray Kranz during one of the most harrowing incidents in the history of human spaceflight. In April 1970, an oxygen tank explosion crippled the Apollo 13 spacecraft roughly 200,000 miles from Earth, turning what was supposed to be the third Apollo landing on the Moon into a desperate fight to bring three astronauts home alive. Kranz was NASA’s lead flight director in Mission Control throughout the multi-day crisis.

What makes Kranz the gold standard for incident command isn’t what he knew. It’s what he did (and more importantly, what he didn’t do).

His job wasn’t to solve the problem

Watch the crisis scenes in the movie closely. Kranz doesn’t try to solve the technical problems himself. He has a room full of brilliant engineers for that. Each of those engineers, in turn, is backed by a “back room” of specialists focused on specific spacecraft systems. Kranz’s job isn’t to solve the problem. His job is to ensure that the problem gets solved, and that’s a subtle but critical difference.

That distinction is fundamental to effective incident command, and many companies get it wrong.

The incident commander role often falls to the most senior engineer by default, with no regard for the practical consequences. The technical work suffers because their best problem-solver keeps getting pulled away to handle logistics and communication, and the overall response suffers because nobody is focused full-time on running it.

What he actually did instead

Kranz was technical enough to understand what his engineers were telling him, and trying to tell each other, but he wasn’t running the calculations himself. So if he wasn’t personally solving the problem, what was he doing instead? His focus was on ensuring that the right people were working the right problems at the right time, not on directly working any of the problems himself.

Watch the movie scenes carefully, and you’ll see a masterclass in incident command:

He organized the response. He set priorities, assigned responsibilities, and made sure every critical task had someone working on it. When the crisis began, he immediately reframed the mission: “From this moment on, we are improvising a new mission: how do we get our people home?”

He demanded clear information. He asked sharp questions and required his team to give him straight answers, not hedged guesses. When the crisis began and readings were contradictory, his instinct was to cut through the speculation: “Let’s work the problem, people. Let’s not make things worse by guessing.”

He made decisions under uncertainty. His engineers disagreed about whether to attempt a direct abort or use the Moon’s gravity to slingshot the crew home. Both options carried enormous risk. Kranz listened to the competing proposals, weighed the tradeoffs, and decided. Then he explained his reasoning so everyone was on the same page.

He kept them looking for possibilities, without being fixated on problems. While his engineers were focused on the individual failures they were trying to address, Kranz stepped back and asked: “What do we have on the spacecraft that’s good?”, challenging them to look for opportunities to use capabilities in unexpected ways to further the mission. When told that a certain system wasn’t designed to do what they needed, he shot back: “I don’t care about what anything was designed to do, I care about what it CAN do.”

He facilitated, not dictated. When engineers brought counter-proposals, he considered them and changed his mind when they were right. When discussion got heated, he reasserted calm: “Let’s hold it down, people.” He allowed debate when it was productive and cut it off when it wasn’t.

He managed the emotional temperature. In military and emergency management circles, this quality is called “command presence.” It doesn’t mean barking orders or being stoic. It means being the steady center of gravity for the response. Responders unconsciously look to the incident commander for cues about how bad things are and whether the situation is under control. If they’re calm and focused, that calmness radiates to the team. If they’re frantic, the team gets frantic too.

Kranz was the calmest person in the room precisely when the room needed calm the most. Not because the situation wasn’t dire, but because he understood that if he lost his composure, everyone else would too.

“Failure is not an option”

Kranz’s famous line from the movie has become a clichรฉ, but it’s worth revisiting. Watch the scene again. It’s not bravado. It’s a statement of intent from someone who has assessed the situation, accepted the constraints, and decided that the team is going to find a way through. He doesn’t pretend the situation isn’t terrible. He tells his team exactly how bad it is, exactly what they need to do, and then sets the expectation that they will find a way to do it.

That’s what good incident command looks like. Not magical thinking, not denial, not heroics. Clear-eyed assessment, clear decisions, and relentless follow-through.

The lesson for your incident response

Your incidents probably don’t involve life-or-death stakes while the whole world watches anxiously. But the model is the same.

The incident commander’s job is to ensure that the problem gets solved, not to solve the problem themselves. That means organizing the response, making decisions, tracking what’s been tried and what hasn’t, communicating with stakeholders, and being the steady center of gravity so everyone else can do their best work.

Many companies default to making their most senior engineer the incident commander. It seems reasonable: they have the most authority, the most experience, the most knowledge.

But your incident commander doesn’t need to be your top engineer. They need to be the person who can stay calm, make decisions with incomplete information, facilitate a room full of smart people who disagree, and ensure that nothing falls through the cracks while everyone else focuses on the technical work.

Gene Kranz showed what incident command looks like when it’s done right. More than half a century later, he’s still the gold standard to aspire to.


I’m writing a book about incident management for software engineering teams, drawing on lessons from both the tech industry and public safety (and movies about NASA!). If you’d like to hear when it’s available, visit im4ds.com.

If your company needs help with incident management right now, my consulting practice is GreatCircle.com/im.

Incident status updates are a translation problem, and the right translator probably isn’t in Engineering

It’s mid-morning in Europe, and your customers are complaining about stale data. It’s 2 AM for your on-call engineers. Your European customer care team noticed the surge in complaints and paged the on-call incident commander (IC). The IC pulled up the dashboards, saw a spike in processing errors on the data ingest pipeline, and paged the data engineering, platform, and networking on-call engineers. Now the IC is coordinating with three responders, all trying to figure out what’s gone wrong, and whether the problem is actually even worse than it appears. And on top of all that, the IC is expected to write a customer-facing status page update.

The result is predictable. The update either doesn’t happen at all, or it reads something like: “Elevated error rates on the ingest pipeline due to Kafka consumer lag.” Which is perfectly accurate, and perfectly useless to 99% of the people reading it.

This is one of the most common patterns I see in my consulting work, and it’s worth understanding why it happens, because the fix isn’t “train your engineers to write better status updates.” The fix is to stop asking your engineers to write them, and start asking the right people instead.

The translation problem

Here’s the core issue: incidents generate information in a language that most of your company doesn’t speak, and most of your customers don’t either.

When responders are working an incident, they talk about services, endpoints, error rates, deployment pipelines, and database replicas. They refer to systems by internal codenames. They describe impact in terms of error budgets and percentile latencies. This is exactly the right level of detail for the people doing the troubleshooting, and exactly the wrong level of detail for almost everyone else.

Your customer support team needs to know which customer-visible features are affected, described in the language customers use. Your sales team needs to know whether the demo environment is impacted and what to say to the prospect they’re meeting in an hour. Your executives need to know the business impact: revenue at risk, customers affected, estimated time to resolution. Your customers need to know that you’re aware of the problem, you’re working on it, and roughly when it will be fixed.

None of these audiences need to know that Kafka consumer lag on the ingest pipeline is causing records to back up in the processing queue. They need to know that some recently submitted data may not be appearing yet, and that your team is on it.

Most companies treat this as a writing-quality problem, but it’s a translation problem. Each of these audiences speaks a different dialect: customer-speak, business-impact-speak, relationship-speak, risk-speak. Your engineers are fluent in engineering-speak, which is precisely what you want from them during an incident. Expecting them to also be fluent in four other dialects, under pressure, at 2 AM, is setting everyone up for disappointment.

Consider how much translation is involved in turning “Kafka consumer lag on the ingest pipeline. Stuck consumer group identified; restarting. Backlog clearing, ETA 30min” into “Some users may notice that recently submitted data isn’t appearing yet. Our team has identified the cause and is deploying a fix. We expect this to be resolved within the next 30 minutes.” Internal service names became customer-visible symptoms. A technical explanation became “identified the cause.” An engineering action became an estimated timeline. That’s the work, and it’s work that customer-facing people do better than engineers.

The right people for the job

The fix is structural. Staff your incident communication functions with people who already speak the audience’s language, and teach them enough incident-speak to do the translation.

A customer support lead who has spent years talking to customers can translate “Kafka consumer lag on the ingest pipeline” into “some recently submitted data may not be appearing yet” far more effectively than an engineer who has never staffed a support queue. An executive liaison who understands how the C-suite thinks about risk can distill a technical sitrep into the three sentences the CEO actually needs, without the incident commander having to figure out what those three sentences are.

Public safety professionals figured this out decades ago. In the Incident Command System used by fire departments and other emergency services, the Public Information Officer (PIO) is a defined incident role with specific training. The principle is straightforward: the people fighting the fire are busy with that, and you want someone else, with special training and experience, speaking to the reporters and the public. The skills are different, the priorities are different, and mixed messaging is dangerous.

What this looks like in practice

The software world’s equivalent to the PIO is an incident role of Communications Lead, or a set of liaison roles, each serving a specific audience.

The simplest version: when a significant incident is declared, someone from your customer-facing team joins the response as a communications liaison. They follow the technical discussion, work from the incident commander’s sitreps, and translate updates into customer-appropriate language for the status page and support team. The IC reviews external communications for accuracy, but the drafting and posting are done by someone who knows how to talk to customers.

This has a second benefit that’s just as important: it frees the incident commander to focus on the response itself. The IC’s attention is one of the scarcest resources during an incident. Every minute they spend crafting a status page update is a minute they’re not spending on coordination, decision-making, or thinking about what they might be missing. Delegating communication beyond the response to someone better suited for the work is better for the IC, better for the response, and better for the customers.

The same principle extends to every other audience. A support update channel where a customer care liaison posts translated, agent-ready talking points. A senior management update path where an executive liaison posts business-impact summaries. A sales notification that flags which key accounts are affected, so reps know before their next customer call. A broad notification to legal, finance, and HR when a significant incident is declared, so each function can self-select whether to engage based on their own expertise, rather than waiting for the IC to figure out whether this incident has regulatory implications.

Three things customers actually want to know

Once you have the right people doing the translating, what they actually need to communicate is simpler than you might expect.

In my experience, what customers truly want to know during an incident comes down to three things: that you’re aware of the problem, that you’re working on it, and, if possible, when it will be fixed. If you cover those three, most customers will be satisfied to let you work the problem in peace.

That’s it. They don’t need the technical details. They don’t need the internal coordination. They don’t need the investigation narrative. All of that is for your team, not for them.

The worst failure isn’t saying the wrong thing

The most damaging communication failure during an incident isn’t saying the wrong thing, it’s saying nothing at all.

A status page that reads “All Systems Operational” while your customers are seeing errors. A support team that has no information to share. Executives finding out about the outage from social media. Each of these erodes trust in a way that even the most jargon-filled status update doesn’t, because silence communicates something very clearly: either you don’t know there’s a problem, or you know and don’t care enough to say anything.

Even an imperfect update is better than no update. Share what you know, be honest about what you don’t, and commit to a cadence for further updates. But calibrate the cadence to the pace of the incident; repetitive content-free “we’re still working on it” updates become noise rather than reassurance. (This is another reason to have skilled communicators handling the job. When I was leading incident management at Slack, the Customer Experience team seemed to have at least a dozen ways of saying “still working on it” without repeating themselves.)

Customers and stakeholders can forgive a rough update. They have a much harder time forgiving radio silence.

This isn’t a problem Engineering can solve alone

Here’s where this gets real: incidents don’t wait for business hours. If your company has 24/7 on-call coverage for engineering, but your customer support and communications teams work 9-to-5, then at 2 AM on a Saturday your engineer is right back to writing status page updates, because there’s nobody else available to do it.

To solve this, the conversation needs to shift from Engineering’s process to the company’s commitment. Staffing incident communication properly means that Support, Customer Success, or whoever owns customer-facing communication needs their own 24/7 on-call rotation, or at least an escalation path that works at 2 AM, even if that means the Director of Customer Care is effectively always on call for high-severity incidents.

It also means those teams need to be included in incident response training (though not necessarily the full responder training that engineering gets; a focused one-hour session on their role during incidents may be all they need). And it means there needs to be a staffing conversation, and probably a budget conversation, with a VP outside of Engineering who hasn’t thought of incident response as their problem before.

That’s uncomfortable, but it’s also clarifying. If your company truly believes that customer communication during incidents matters, then the teams who are best at customer communication need to be available when incidents happen. If those teams aren’t willing to staff for it, that tells you something about how seriously the company takes it, which is useful information in itself.

For the engineering leader reading this: you shouldn’t try to fix this alone (and you probably can’t anyway). What you can do is make the case. The argument is straightforward: Engineering has invested in 24/7 incident response capability because incidents don’t respect business hours. Customer communication during those incidents is at least as important as the technical response, and it requires a different skill set. The same logic that justifies an engineering on-call rotation justifies a communication on-call rotation.

If the company isn’t willing to make that investment, then it’s making a conscious choice to accept poor customer communication during off-hours incidents, and everyone should be honest about that tradeoff.

The real question is organizational design

The question to ask isn’t “how do we train our engineers to write better status page updates?” The question is “who in our company already has the skills to communicate with each of these audiences, and how do we bring them into our incident processes?”

That’s an organizational design question, not a training question. And it’s the kind of question that separates companies that handle incidents well from companies that just handle the technical parts well.

Remember that European customer care team from the opening? They were the ones who noticed the problem. They were the ones fielding the complaints. They already knew how to talk to the affected customers. All they needed was a seat at the table, so they could translate what the engineering team was seeing. That’s not a big ask. But it’s one that most companies have never thought to make.


I’m writing a book about incident management for software organizations, drawing on my experience building and leading incident management programs at companies like Google, Slack, and others. If you’d like to hear about it when it’s available, visit im4ds.com.

If your organization is wrestling with incident communication or other incident management challenges, I do consulting and training on these topics. You can find out more at greatcircle.com.

In incidents, swarming is a feature, not a bug

Spontaneous swarming of responders might seem like a nuisance that breaks our tidy mental models of incident response, but it’s actually very powerful. It’s something to facilitate and encourage, not simply tolerate.

Most of us expect incident response to be fairly linear: an on-call engineer gets paged, decides an incident is called for, and pages an incident commander. The IC spins up an incident channel and pages a couple of additional responders. So far, so good.

What surprises many people is that, within minutes, a bunch of additional reinforcements often begin swarming the incident, without having been paged. A leader at one company I worked with described this as their “white blood cell response”: something goes wrong, and people spontaneously converge to help however they can.

Why do these volunteers show up? For different reasons. Some are engineers who recognize the symptoms from something they’ve seen before, or who know that their systems might be involved. Some are from customer care, account management, or product, keeping an eye on the incident to see if it’s going to affect their areas of concern.

Many of them aren’t there to actively investigate; they’re monitoring, ready to speak up if they have something to contribute. This is one of the advantages of running incidents in Slack rather than on a Zoom call: people can follow along and speak up without disrupting the response. They don’t have to wait for a pause in conversation or work up the courage to interrupt, which in my experience causes many useful observations to go unsaid.

What these spontaneous responders all have in common is that they’re busy experts who have decided that the incident is worth their attention. That’s an organizational strength worth recognizing, developing, and leveraging.

So how do you facilitate and encourage swarming? Start by understanding how incidents actually unfold.

The model vs. the reality

Most incident management processes are built around a tidy model: something happens, an incident is declared, the incident commander pages the right people, those people respond, and the organized response proceeds from there. It’s a useful model, and it describes the formal structure of incident response well enough, but like all models, it has its limits.

Real incidents are messier than that. People find out about problems in all kinds of ways. Engineers start investigating before an incident is even declared. Volunteers show up before the IC has finished their first round of pages. By the time the formal response is organized, a significant chunk of the actual response is already underway, whether it’s coordinated or not.

Often the swarm forms without any incident tooling involved at all. An engineer investigating a problem asks a coworker for help. That coworker pulls in another. Before long, there’s an ad hoc team working on an urgent, significant problem in whatever random place the discussion started, usually some team’s day-to-day channel, where it’s intermingled with “where are we going to lunch today?” and bot posts about new code PRs and deploys. Eventually someone asks “hey, maybe we should be treating this as an incident?” (Pro tip: the answer to that question is almost always an emphatic “yes!”)

When I was leading incident management at Slack, we maintained a standing channel called #incident-next, ready for whoever needed it. Anyone investigating a concern could direct people there to start collaborating in a focused space, separate from whatever team channel the discussion started in. When an incident was declared, the channel got renamed to become the incident channel, preserving all the context, the investigation history, and the people who were already there, and a new empty #incident-next was created to be ready for the next one. This is a simple mechanism that any company could implement today, and it directly supports swarming: people converge, start investigating, and the formal incident declaration catches up to what’s already happening. The response is already underway when the structure arrives to support and harness it.

Most companies design their process only for the tidy model and then get frustrated when reality doesn’t match. A better approach is to design for how incidents actually unfold, which means designing with swarming in mind as a possibility, or even something to be encouraged.

Why swarming works

In a purely command-driven response model, someone has to figure out which teams are relevant, page the right on-call engineers, wait for them to respond, brief them on the situation, and assign them work. Whether that’s the incident commander or a tech lead, every step is bottlenecked through one person’s knowledge and judgment. If they page the networking team but the problem turns out to be in the caching layer, time is lost while they regroup and page a different team. They can only route expertise as well as they can diagnose the problem, and in the early minutes of an incident, the diagnosis is often wrong.

Swarming bypasses that bottleneck. The people with relevant knowledge self-select in, often before the IC or TL even knows their expertise is needed. The engineer who made a related change earlier that day, the one who saw a similar failure pattern last month, the one who happens to know that the caching layer was acting strangely yesterday, they show up because they can see that something is wrong and they recognize that they have something to contribute. Nobody leading the response could have paged all the right people that quickly, because nobody could have known who the right people were.

On a small incident, this is often all you need. A handful of knowledgeable people converge, naturally divide the work, and the problem gets solved quickly, sometimes faster than a formal response could even get organized.

On larger incidents, swarming gets the right people into the room faster than any paging process could. The IC still coordinates the response, but instead of spending the first forty minutes figuring out who to call and waiting for them to respond, they’re organizing people who are already there and ready to work. That’s a fundamentally better starting position. The trick is to make sure that the room is ready for the swarm.

When it isn’t, the same instinct that drives swarming produces freelancing instead: people investigating on their own, without coordinating, because there was no organized response to join. Same motivated people, very different outcome.

Facilitating the swarm

If swarming is this effective, why doesn’t it work well at every company? First, people have to feel safe showing up uninvited. In a culture where volunteering in an incident means risking blame if things go sideways, the swarming instinct gets suppressed before it starts. That’s a deeper problem than process design can solve, but it’s worth naming.

Assuming the culture supports it, two practical things determine whether swarming works:

Can people find the response? When someone hears about a problem, whether it’s a teammate mentioning that checkout is broken, an alert in their team’s channel, an error spike on a dashboard, do they also hear that there’s an organized response underway?

Maybe this is an #all-incidents channel where every declared incident is automatically announced with a link to the incident channel. Active incidents surfaced on the dashboards people are already checking. A predictable naming convention for incident channels, so that anyone can type #inc- in Slack and use auto-complete to see what’s active right now. Integration with your chat platform so that people who are talking about the problem in their team channels get a pointer to the organized response.

The goal is to make sure that everyone who hears about the problem, however they hear about it, also learns where people are converging.

Then, can they easily contribute? Once people find the channel where the response is happening, they need a clear on-ramp. A pinned message with the current situation, who’s filling what role, and what’s being investigated. A clear expectation that new arrivals announce themselves and what expertise they bring, then self-brief by reading the channel and the latest situation report before asking for an assignment. Enough structure that someone can get oriented in a few minutes and start contributing, rather than facing a wall of scrollback with no obvious way to plug in.

None of this is complicated. But it does require deliberate design. The swarming instinct happens naturally. The infrastructure that makes swarming productive doesn’t.

Invest in the instinct

If your team is already swarming incidents, you’ve got something invaluable: people who care enough to show up. The rest is infrastructure: making the response visible, making it easy to join, and giving the IC the tools to organize a growing swarm on the fly.

You can’t manufacture that instinct. But you can make it dramatically more effective.

Swarming is one of the topics I cover in depth in my forthcoming book, Incident Management for DevOps and SRE. If you’d like to know when it’s available, sign up at im4ds.com. And if your company needs help with incident management right now, my consulting practice is greatcircle.com/im.

Most Companies Wait Too Long to Declare Incidents

Most companies have some notion of what an “incident” is: a significant problem that requires an urgent response and involves multiple responders. But having a definition and actually using it are two different things. Teams at most companies I’ve worked with know perfectly well what an incident is, at least in theory. In practice, they still hesitate to declare one, for reasons that have nothing to do with the definition.

The pattern is remarkably consistent. An engineer sees something that looks wrong and thinks “maybe it’s not that bad.” A team notices a problem that could be an incident, but nobody wants to “bother” the on-call incident commander over what might turn out to be nothing. Someone suspects they should declare, but waits for a more senior person to make the call, assuming it’s not their place. A customer care agent gets their third chat this hour about the same issue, and wonders if it’s just a coincidence. An account exec gets a curious “are you folks having any problems today?” from their contact at a major customer, but isn’t sure whether or how to escalate. Across the company, people are smelling a whiff of smoke, but nobody is pulling the fire alarm.

When I assessed one company’s incident management program, stakeholders estimated they were capturing only 75-80% of actual incidents. The rest were being handled informally, in ad hoc threads and side conversations, without the coordination, communication, and documentation that the incident process provides. That’s typical, in my experience; if anything, 75-80% is better than most.

The interesting question isn’t whether your team waits too long, or avoids declaring at all. They almost certainly do. The interesting question is why.

It’s Not a Judgment Problem

The obvious explanation is that people don’t know where the line is. If only we had clearer criteria, the thinking goes, people would declare at the right time. So companies invest in decision trees and flowcharts and severity matrices, but the problem doesn’t get better.

It doesn’t get better because unclear criteria aren’t the real issue. The real issue is that your company has made declaring an incident costly and risky for the person who does it.

Think about what happens when someone declares an incident at your company. I’ll bet it looks something like this:

They’re committing real time and attention, and not just for themselves. They’ve just pulled themselves and several colleagues away from whatever they’d planned to be working on, and added several hours of incident-related work to each of their already-overflowing plates.

There’s no lightweight way to raise an urgent concern. There’s no easy way to say “I think something is wrong right now” and get someone experienced to look into it quickly. Filing a bug or ticket puts the problem in a queue; the only way to get an immediate response is to formally declare an incident, with all the overhead that entails. The person who notices the problem has to decide that it’s an incident, assess how bad it is, and figure out who to page. That’s a lot to expect from whoever randomly happens to notice first, and almost guarantees that nothing happens until somebody senior enough or confident enough eventually notices the problem.

They’re triggering a heavyweight process. Declaring the incident guarantees that a heavyweight post-incident review process kicks in, with mandatory documentation and required meetings. The incident gets counted, and someone in leadership is tracking that count, wanting it to trend downward.

They’re paying a career cost. The time engineers spend responding doesn’t earn them any credit on their performance review; at best it’s invisible, and at worst it’s counted against the “real” work they didn’t finish. Their sprint commitments don’t get adjusted. Their manager doesn’t say “I see you spent 15% of this quarter helping with incidents, so let’s recalibrate your goals.” The incident work just disappears into an unacknowledged gap between what they delivered and what was expected.

Engineers aren’t oblivious to these incentives, even if they couldn’t name them. They’re responding to them rationally. When declaring an incident is costly, for themselves and for the coworkers they’d be pulling in, people unconsciously raise their internal threshold for what’s worth declaring. They wait a little longer, hoping the problem resolves itself. They try to fix it quietly before anyone notices. They let someone else make the call.

Declaring Early Beats Declaring Late

Here’s what companies often overlook: the costs of declaring too early and declaring too late are not the same.

If someone declares an incident and it turns out to be a false alarm, the cost is small. A few people spend a few minutes getting oriented, realize the situation is under control, and stand down. In fact, it’s not really a cost at all; your team just got a bit of practice with the incident process, which is valuable in its own right.

If someone doesn’t declare an incident and the problem turns out to be serious, the costs are large. Customers are affected longer. The blast radius expands. What could have been a contained, quickly resolved incident becomes a prolonged outage. And the longer a problem goes unaddressed, the harder it gets to recover from.

One of my flight instructors taught me a lesson that I apply constantly in incident management: “If you wonder whether you’re running out of fuel, you’re running out of fuel.” If your subconscious is even raising the question, it’s telling you something. The same principle applies here. If someone is wondering “should we be treating this as an incident?”, the answer is almost certainly “yes!”

Fix the Incentives

If your team is slow to declare, the fix usually isn’t better decision flowcharts. It’s examining the incentives, both explicit and implicit, that your company has built around incident declaration, and changing the ones that inadvertently punish people for doing the right thing. Here’s where to start.

Make declaring cheap. Not every incident needs a full post-incident review. Lightweight incidents that resolve quickly should have a lightweight process. If every declaration triggers the same heavyweight machinery regardless, people will avoid declaring in order to avoid the machinery.

Separate reporting from declaring. When someone sees something that might be an incident, they should be able to flag it quickly, without having to determine the severity, identify which teams to page, or commit to a formal declaration.

Think of it like calling your local emergency number (911, 999, 112, or whatever your country’s is): the caller describes what they’re seeing, and the person who answers (who has had special training and lots of experience in evaluating reports of possible emergencies) decides what response is called for. The person who reports the problem doesn’t bear the weight of all those “is it an incident? how severe? who do we page?” decisions. They just raise the alarm, and an expert (typically an on-call incident commander) takes it from there.

Stop fixating on incident counts. When leadership treats the number of incidents as a metric to drive down, the predictable result is that people stop declaring. The number of incidents you declare should not be a target. In fact, a rising count can be a healthy sign: it may mean people are becoming more comfortable with the process and using it more. Signs of success such as more customers, more features, more employees, and more usage can all lead to higher incident counts.

Recognize incident work as real work. Time spent responding to incidents needs to be visible in performance ratings, bonuses, and promotion decisions. Not as a footnote, not as an “also did,” but as a genuine contribution to the company. If the only work that counts is feature delivery, and working on incidents is a distraction from that, then your best engineers will subconsciously but rationally avoid incident work.

Watch for mixed messages. Companies often undermine their own stated values without realizing it. A leader pressures a team about incident counts while simultaneously asking why problems aren’t caught earlier. A manager expresses frustration about “unnecessary” declarations while wondering why the team doesn’t escalate fast enough. A VP asks why incidents aren’t caught sooner, then questions why their team is spending so much time on incidents. The contradictions send a clear message about what’s actually valued, regardless of what’s written in the incident management policy.

These mixed messages aren’t limited to Engineering. Noticing and resolving these contradictions, at every level and throughout the entire company, is what makes timely declaration a cultural reality instead of just an aspiration.

When in Doubt, Declare

The next time someone is wondering “is this bad enough to declare?”, you want “yes!” to be an easy call. Not because you’ve written better criteria, but because you’ve built a system where the cost of declaring is low, the process is lightweight, and the culture rewards raising the alarm.

All your incentives and processes should align to reinforce one simple principle, without fear or reservation:

When in doubt, declare!


This is one of the foundational questions I tackle in my book, Incident Management for DevOps and SRE, which I’m currently writing. If you’d like to be notified when it’s available, sign up at im4ds.com.

If your company is wrestling with these questions right now and doesn’t want to wait for the book, my consulting practice can help.

What Is an “Incident”?

Ask six people at the same company what counts as an incident, and you’ll likely get six different answers.

The IT ops manager thinks of every ticket in ServiceNow. The on-call SRE thinks of the alert that paged at 2 AM. The head of sales thinks of the call from an angry enterprise customer. The finance team thinks of the disruptions they have to issue SLA credits for. The engineering manager thinks of last week’s outage, the one she was up until 3 AM for, acting as the incident commander and coordinating the response across three teams. The newest engineer isn’t sure, but knows they don’t want to be the one who declares one.

This confusion isn’t academic. When people in the same organization mean different things by “incident,” practical problems follow. The CFO asks “how many incidents did we have last quarter?” to forecast SLA payouts for the board report, and gets a completely different number depending on who answers. Engineering can’t tell whether the trend line is improving because the data mixes SRE pages, ITIL tickets, and multi-team coordinated responses into a single count. Someone declares an incident and half the organization thinks it’s a crisis while the other half thinks it’s Tuesday. And every attempt to improve incident management stalls because the people in the room haven’t realized they’re talking about different things.

Where the Confusion Comes From

The word “incident” has been overloaded with several overlapping meanings in the technology industry; people think they’re all talking about the same thing, but they aren’t, and the confusion causes more problems than most people realize.

ITIL has taught generations of IT professionals that an incident is any unplanned interruption to a service. Under that definition, a single user unable to log in is an incident. A slow database query is an incident. A printer jam is an incident. ITIL-trained teams sometimes process thousands of “incidents” per month through their ticketing systems. When someone from that background hears “we need better incident management,” they’re thinking about ticket queues and resolution targets, not about coordinated emergency response.

PagerDuty, the most widely-used on-call tool in the industry, has long compounded the problem through its product terminology. In PagerDuty’s data model, which stretches back over 15 years to the early days of the company, every alert that pages someone creates an “incident.” You literally cannot page a colleague without creating a PagerDuty “incident.” This means an on-call engineer who gets paged three times on a quiet Tuesday has, in PagerDuty’s language, experienced three incidents. Over time, this trains people to think of “incident” as synonymous with “page” or “alert.” PagerDuty has recognized the gap as they’ve expanded into incident response orchestration, particularly after acquiring Jeli in 2023, adding the concept of “major incidents” (a term ITIL also uses, for the same reason) to distinguish coordinated-response situations from everyday pages. But the underlying product model remains, and generations of engineers have already internalized the equation: incident = someone got paged.

Then there’s the other extreme. Some organizations reserve “incident” exclusively for the worst events they can imagine: full-site outages, data breaches, events that make the news. This sounds disciplined, but it creates a high psychological barrier to declaration. If “incident” means “catastrophe,” nobody wants to be the person who declares one for something that turns out to be just a blip. So people hesitate, they wait for clearer signals before acting, and by the time someone finally says the word “incident,” the situation has been burning unchecked for longer than it needed to.

Those are all legitimate uses of the word, but they’re answering different questions. The ITIL definition tells you what to put in the ticketing system. PagerDuty’s tells you what to attach a page to. Finance’s tells you what catastrophes will trigger an SLA payout. None of them help you decide when to urgently pull people together and mount a coordinated response.

We’re not saying those other definitions are wrong and ours is the One True Definition. We just want folks to be aware of the potential for confusion, and to be sure they understand from context (or explicit clarification) which version of “incident” someone is referring to.

It’s like the word “security,” which means something completely different to an information security engineer, a physical security guard, and a financial analyst trading stocks and bonds. Nobody argues about which meaning is correct; they just make sure everyone in the conversation knows which one they’re talking about.

My professional community, my own work, and my forthcoming book are about managing emergencies that require urgent, multi-person, coordinated responses. In those contexts, we use “incident” to mean a situation that is significant enough to need attention right now, urgent enough that you can’t just file a ticket and walk away, and beyond what one person can handle alone. That’s one valid, context-specific definition among several. The important thing is that everyone in a conversation is clear on which one they’re using in the moment.

Watch for It

Here’s a concrete step you can take this week: ask six people at your company what they think counts as an incident. Don’t ask in the abstract; give them three or four scenarios and ask which ones they’d call an incident. Better yet, do it in a group setting where everyone can hear each other’s answers. The spread will surprise people, and that surprise is the point.

Once your team sees the mismatch, you gain the ability to catch it in the moment. “Wait, are we all talking about the same kind of incident right now?” is a surprisingly useful sentence.


This is one of the communication challenges I explore in my book, Incident Management for DevOps and SRE, which I’m currently writing. If you’d like to be notified when it’s available, sign up at im4ds.com.

If your organization is wrestling with these questions right now and doesn’t want to wait for the book, my consulting practice can help.

Self-briefing: how to join an incident without interrupting to ask “what’s going on?”

It happens on way too many incidents. The ad hoc team of responders is deep in discussion when somebody new joins the call or the channel, and the first thing they say is: “Hey, what’s going on?”

The incident commander (IC) stops what they’re doing and recaps the timeline, the current theories, who’s working on what, and what’s been tried so far, which can easily take a few minutes. Meanwhile, the discussion and coordination stall, other responders wait (or, worse, proceed without coordination), and the IC loses their train of thought. Then fifteen minutes later someone else joins, and it happens again.

Each interruption seems small. Cumulatively, they’re one of the biggest drags on incident response that nobody talks about. And it’s not just responders. Executives, customer care reps, sales teams, and other observers who join the channel to follow along often do the same thing, without realizing how much these interruptions add up.

The problem isn’t the people

People who ask “what’s going on?” aren’t being lazy or inconsiderate. They’re doing what feels natural: they’ve been pulled into an unfamiliar situation, they don’t know what’s happening, and they turn to the person in charge for orientation. It’s a reasonable instinct.

But it’s also an expensive one. The incident commander is typically the busiest person in the response. They’re tracking multiple threads of investigation, coordinating assignments, communicating with stakeholders, and maintaining the big picture. Every time they stop to deliver a verbal briefing, all of that pauses. And because each new arrival gets a slightly different recap depending on when they ask and what the IC remembers to mention, the team can end up with inconsistent pictures of the situation.

A new arrival who doesn’t know what’s going on has three options, none of them good. They can interrupt the IC for a recap, which is disruptive. They can wait until someone has time to bring them up to speed, which delays whatever they’re there to do. Or they can start acting on incomplete information.

There’s a better alternative: self-briefing. It’s one of the simplest, highest-impact changes a team can make to their incident response, though it does require laying some groundwork.

What self-briefing looks like

Self-briefing means that when you join an incident in progress, you orient yourself rather than interrupting to ask someone else to orient you. For observers following along from the exec team or customer care, this can be completely silent: read the channel, read the latest situation report (sitrep), and you’re up to speed without anyone even knowing you arrived. For responders, a two-message pattern lets other responders know you’re there and coming up to speed, with minimal disruption.

Message one, on arrival: “Hi, I’m Jordan from the caching team. Reading back through the channel now.”

This tells the IC and the rest of the team several important things: you’re here, you have potentially relevant expertise, and you’re getting yourself up to speed rather than asking them to stop and brief you.

Then you actually do the reading. Scroll back through the incident channel. Read the most recent sitrep, if one has been posted. Check the status document if there is one. Note who’s involved and what they’re working on. This typically only takes a few minutes, depending on how long the incident has been running and how much has happened.

Message two, when you’ve briefed yourself: “Jordan from caching, caught up. I’ve read the channel and the latest sitrep. It looks like nobody is currently investigating the CDN layer. Want me to start looking there, or is there something else that would be more useful?”

Now your first real interaction with the IC is useful rather than disruptive. You’ve demonstrated that you understand the current situation. You’ve identified a potential gap. And you’ve offered a specific proposal for what you could work on, which is much easier for a busy IC to respond to than an open-ended “what do you need?”

Why this only works with text channels

Self-briefing depends on text channels. It works well in Slack (or a similar channel-oriented text tool), but isn’t feasible in Zoom (or any voice/video-first approach).

In Slack, the incident channel gives a new arrival something powerful: control over their own depth. You can skim the last twenty messages for a quick sense of where things stand, then slow down and dig into the thread about cache hit rates because that’s your area and the detail matters to you specifically. You can click the dashboard link someone shared and look at the graph yourself. You can see who said what, which tells you who’s working on what and who to direct your follow-up questions to. You can reread a confusing message until it clicks. You can forward a specific message to a colleague and ask “what does this mean?” or “this seems critical for your team.”

With Slack, a few minutes of reading lets you go shallow for orientation and deep where your expertise is relevant, all in the same pass.

In a Zoom call, that context evaporates the moment the words are spoken. You can’t rewind a conversation. When a new responder joins a call that’s been running for forty-five minutes, those forty-five minutes of discussion are gone. And it’s not just the words. Every graph someone screen-shared, every dashboard someone walked through, every log snippet someone pasted into the Zoom chat (which late joiners also can’t see) is gone too.

With Zoom, the only option is exactly what we’re trying to avoid: asking someone to stop and recap.

“But what about transcripts?” Even if your video conferencing tool offers AI-generated “catch me up” summaries for late joiners, those give you the shallow orientation pass and nothing else. They can’t show you the graphs and dashboards people were discussing while screen-sharing. They struggle with the names of people, systems, and services that fill every incident conversation. And they don’t let you drill into the specific thread that matters most to a person with your particular expertise. An AI summary might tell you the team is investigating a caching issue. It won’t let you read the actual exchange between the two engineers who narrowed it down, click through to the dashboard they were looking at, and decide for yourself whether they’ve checked the CDN layer yet.

Some teams start out with the best of intentions, keeping the Slack channel updated alongside the Zoom call. In practice, it works for a little while, and then the person doing the updating gets drawn into the verbal conversation and stops scribing. Critical information, discoveries, and decisions never make it into text. The Slack channel becomes a sparse, incomplete shadow of what’s actually happening on the call.

This isn’t a minor inconvenience. It’s a structural barrier to self-briefing. If your primary incident communication happens on a voice call, you’ve made it physically impossible for new arrivals to orient themselves without interrupting someone. You’ve baked the “what’s going on?” problem into your communication architecture.

Periodic situation reports posted on a regular cadence by the incident commander help bridge this gap, because a good sitrep gives a new arrival a snapshot of the current state regardless of what communication tool the team is using. But sitreps are periodic summaries. They can’t let you explore the details that matter for your specific expertise. A team that communicates primarily in text gets both: the running record you can explore at your own depth in the channel, and the periodic snapshot in the sitrep.

Make it an organizational expectation

Self-briefing sounds simple, and it is. But it won’t happen consistently unless the organization establishes it as an explicit expectation, not just a nice idea. This means a few things.

Teach the pattern. Include self-briefing in your incident response training. Teach responders the two-message template. Explain why it matters. Most people will adopt it readily once they understand the reasoning; they just need to know it’s expected.

Lay the groundwork. Self-briefing depends on having something to brief yourself from. That means using a text channel as your primary communication tool during incidents, posting regular sitreps, and maintaining a status document on longer incidents. If your team’s incident communication happens mainly on a Zoom call with a neglected Slack channel on the side, there’s nothing for a new arrival to self-brief from. The expectation and the groundwork go hand in hand: each one reinforces the other.

Reinforce it from the IC role. When a new responder joins and immediately asks “what’s going on?”, the IC can gently redirect: “Welcome! Take a few minutes to read back through the channel and check the latest sitrep, then let me know what questions you have.” A few consistent redirections establish the norm quickly.

Model it yourself. When you join an incident as a responder, follow the two-message pattern even if you could get a faster verbal briefing. Especially if you’re senior. When a staff engineer or a VP visibly self-briefs rather than expecting a personal recap, it sends a powerful signal about how things work here.

The compounding benefit

Self-briefing doesn’t just help the person who arrives prepared. It helps everyone.

The IC stays focused on managing the response instead of delivering repeated briefings. The investigation maintains momentum because nobody is hitting pause to catch up new arrivals. The existing responders don’t lose context from their own work while waiting for the IC to finish briefing someone else. Observers from customer care or the exec team can follow along and update their own stakeholders without pulling anyone away from the response. And the new responder starts contributing faster, because five minutes of reading usually builds better context than a rushed verbal summary anyway.

Over the course of a large incident where a dozen people join at different times, the difference between “everyone self-briefs” and “everyone asks the IC for a recap” can easily be an hour of the IC’s time, and several hours of cumulative disruption to the response.

The technique is simple, but the impact is significant. For observers, self-briefing is invisible: just read the channel and the latest sitrep. For responders, it takes two messages and a few minutes of reading. Nobody gets interrupted. Nothing stalls. And the response keeps its momentum.


Self-briefing is one of the practices I cover in my forthcoming book, “Incident Management for DevOps and SRE.” If you’d like to hear when it’s available, sign up at im4ds.com. And if your organization needs help building effective incident management practices right now, my consulting practice is greatcircle.com/im.

Is your team reacting to incidents, or responding?

Think about what happens when a fire alarm goes off in a hotel.

The guests are jolted awake at 2 AM, groggy and disoriented, by a blaring alarm in an unfamiliar room. They fumble for their shoes and coats, grab their phones (but forget their room key), then try to find the exits through hallways they’ve only seen once, hours ago when they checked in. The elevators are disabled, so they stumble down 14 flights of stairs with a crowd of other half-awake, disgruntled guests. They gather outside and wait for someone to come and tell them whether it’s safe to go back in.

For the guests, this is a disruption at best and a crisis at worst. Their night has been upended by something they didn’t expect and can’t control, and they’re wondering how much sleep they’ll get before their big meeting tomorrow. All they can do is stand outside in the cold, and hope someone else fixes it soon.

The hotel staff do what they can: directing guests toward exits, calling 911, meeting the fire department at the entrance. But they’re a handful of people with limited training managing a building full of guests who are confused, annoyed, and frightened.

The guests and staff react.

Now think about what happens when the fire department arrives.

The firefighters respond.

Before the first crew even steps off the fire engine, their officer gets on the radio: “Engine 4 on scene, nothing showing, investigating.” Then they check the alarm panel, do a size-up, and start working through well-practiced procedures they’ve followed so many times that they’re second nature. More units are on their way, and even more are just a radio call away, if needed. For the fire department, this isn’t a crisis. It’s just another call in a routine shift.

The difference between reacting and responding isn’t about who cares more. The fire department cares deeply about the safety of the people in that building. The difference is preparation. The guests have no plan, no training, no tools for this situation; they can only react. The fire department has all of those things; they can respond.

The same pattern shows up in software incidents

When something breaks at 2 AM and your on-call engineer gets paged, what happens next? Do they poke at the problem alone, hoping they can fix it before anyone notices? Do they post a vague message in Slack, and then three different people start digging into the same thing without coordinating with each other? Does a senior leader show up and start barking orders, whether or not they have context?

That’s reacting; it’s what happens when people encounter a problem they haven’t prepared for. It’s the natural result of not having a plan, roles, and practiced procedures in place.

Responding looks different. Someone pages an incident commander (IC), and an incident gets declared. The IC assesses the situation and sets initial priorities. Responders are assigned to specific tasks. Communication flows through well-understood channels. Status updates go out at regular intervals. People know what their role is, what’s expected of them, and how to work together effectively in an emergency.

Most engineering teams are full of smart, committed people. What separates chaos from a coordinated response is preparation.

A diagnostic question for your organization

This distinction is one of the most useful questions you can ask about your organization’s incident management maturity: when something goes wrong, does your team react, or respond?

Here are some signs you’re still reacting:

There’s no clear moment when “normal work” shifts to “incident response.” People gradually realize something is wrong and start working on it individually, without explicit coordination.

There’s confusion about who’s in charge.

Multiple people investigate the same thing without knowing it.

Status updates happen sporadically, if at all.

Senior leaders don’t know what’s happening and start asking questions that pull responders away from the work.

When it’s over, nobody is quite sure whether or when to stand down.

Responding, by contrast, has clear transitions: a declaration that shifts the team into a different operating mode, defined roles that people step into, communication practices that keep everyone informed, and an explicit close-out that tells people the emergency is over and they can return to their regular work.

The shift from reacting to responding is incremental, not instant

Most organizations start out reacting. That’s natural. You can’t respond to something you haven’t prepared for, and most organizations don’t invest in incident management preparation until they’ve been burned by a few incidents that didn’t go so well.

The good news is that you don’t need to build all of this overnight. Start with the basics: a clear way to declare that an incident is happening, someone designated as the IC, and a shared communication channel for the incident. That alone will move you from pure reaction toward coordinated response. Then build from there, adding structure, process, and tooling as your team gets comfortable with each new piece.

Your first few formally managed incidents will feel awkward and clunky. That’s fine. The fire department’s recruits feel that way on their first calls too. What matters is that you’re building the capability, one incident at a time. Every incident you manage with even a basic structure is a repetition that makes the next one smoother.

The goal isn’t perfection. It’s preparation. Because when the fire alarm goes off (and it will), the question is whether your team is prepared to respond instead of react.

I’m writing a book about building these capabilities: Incident Management for DevOps and SRE, a practitioner’s guide to structured, effective incident response. If you’d like to hear when it’s available, sign up at im4ds.com.

And if your organization needs immediate help building these capabilities, well, that’s what my consulting practice at Great Circle is all about.

The social contract of emergency mode

When a fire engine rolls down the street without its lights and sirens on, it’s just a big red truck with a fancy paint job. It follows the same traffic rules as any other vehicle its size. Almost nobody pays it any special attention.

But the moment the crew gets dispatched to a call and flips on the lights and sirens, the rules change, not just for the firefighters, but for everyone around them. Other drivers pull over, and cross traffic yields at intersections (supposedly, anyway). Everyone understands that a different set of rules is now in effect, and that those rules are temporary. When the lights and sirens shut off, normal rules resume.

Software organizations need the same kind of shift when an incident happens, but we have to create the signal ourselves, because we don’t have lights and sirens (most of us, anyway; ask me sometime about driving Code 3 and going double the speed limit… at Burning Man, where the speed limit is 5 miles per hour). That signal is the explicit declaration of an incident.

Two modes, one organization

Most technology organizations operate day-to-day in what I think of as “normal mode.” Decisions are made through discussion, deliberation, and consensus-building. Organizational structure follows reporting relationships and seniority. Time is measured in weeks, months, and quarters. This is the right way to run a software organization most of the time.

But when customers are being impacted by an outage, normal mode doesn’t work. You can’t spend three days building consensus on whether to roll back a bad deployment while your service is down. You can’t route a decision through two layers of management while errors are piling up.

When you declare an incident, you shift into “emergency mode.” Time gets measured in minutes and hours. A temporary organizational structure takes effect, where an incident commander (IC) is in charge regardless of where anyone sits in the everyday org chart. A mid-level engineer serving as IC might be coordinating the work of senior engineers and directors. That’s not a problem; it’s the design working as intended. And decision-making becomes more directive: the incident commander makes decisions after considering input, but doesn’t wait for perfect consensus.

This shift feels uncomfortable if you work in a culture that values flat structures and consensus-building. Good. It should feel uncomfortable as an everyday way of operating. Emergency mode isn’t a better way to run an organization. It’s less inclusive, less thoughtful, and more prone to blind spots. But decades of experience in public safety and other fields that deal with emergencies have shown it’s the most effective way to get through an emergency, so you can return to your normal, more collaborative way of working as quickly as possible.

Turning the lights and sirens on (and off)

Here’s the part many organizations get wrong: they never make the shift explicit.

Without a clear declaration that an incident is underway, you get mismatched expectations. Some people treat the situation as an emergency while others respond at their leisure. Some people feel intense urgency while others don’t realize there’s a problem at all. The response becomes uncoordinated, not because people don’t care, but because they’re operating under different assumptions.

Declaring an incident is how you turn on the lights and sirens. It tells everyone: different rules are now in effect. Expect faster communication. Expect a temporary organizational structure. Expect more directive decision-making. This is not how we normally operate, and that’s intentional.

But declaring the end of an incident is just as important as declaring the start. Without an explicit “all clear,” responses fizzle out instead of ending cleanly. People aren’t sure whether they’re still expected to be available at incident-level speed. The on-call engineer who was paged at 2 AM doesn’t know whether they can actually go back to sleep. The incident channel stays open for days with low-priority chatter that nobody feels empowered to shut down.

The IC should explicitly close out the incident: acknowledge everyone’s contributions, confirm the service is restored, and point people toward the post-incident review. This gives responders permission to disengage and return to their normal work. It creates a clear boundary between emergency and normal mode. And it preserves emergency mode as something meaningful, not just a more chaotic version of everyday operations.

The contract

This shift between modes is, at its core, a social contract. You’re asking people to operate differently for a while: to make faster decisions with less information, to set aside some of their usual ways of working, to accept more directive leadership. In return, you’re promising that this is temporary, that it’s only happening because it’s genuinely necessary, and that you’ll return to normal as soon as you can.

The contract breaks down in predictable ways. If leadership declares incidents for things that aren’t really emergencies, people stop taking incident declarations seriously. If responders refuse to shift their behavior during actual emergencies, insisting on the same level of debate they’d use for a design review, incidents drag on unnecessarily. If ICs don’t declare when incidents are over, people get stuck in emergency mode and burn out.

And there’s a particularly damaging failure mode: organizations that are always in emergency mode. If everything is an emergency, nothing is. People become numb to the urgency. They stop responding with the focus and intensity that real emergencies demand. The social contract erodes, and when a genuine crisis hits, the organization discovers it’s lost the ability to shift gears.

Making it work

The most capable organizations I’ve worked with treat this mode-shifting as a core discipline. They declare incidents explicitly. They staff defined incident roles, like incident commander, from a trained pool of people who have other jobs most of the time. They operate under emergency rules for as long as needed, and not a minute longer. And they return to normal deliberately, not by just letting things wind down.

This isn’t about having perfect processes or expensive tooling. It’s about your organization agreeing, in advance, on what emergency mode looks like, when to invoke it, and how to exit it. It’s about building the muscle memory to shift gears when customers need you to, and the discipline to shift back when the emergency is over.

The ability to make this shift cleanly and confidently is one of the clearest markers of incident management maturity I’ve seen across the organizations I’ve worked with. Getting it right changes how your team experiences incidents: from chaotic and draining, to intense but manageable.

This is one of the ideas I’m developing in my forthcoming book, Incident Management for DevOps and SRE. Sign up at im4ds.com to be notified when the book is available, and to receive occasional progress updates and early access to selected content. If your organization needs help with incident management right now, my consulting practice is GreatCircle.com/im.