Is your team reacting to incidents, or responding?

Think about what happens when a fire alarm goes off in a hotel.

The guests are jolted awake at 2 AM, groggy and disoriented, by a blaring alarm in an unfamiliar room. They fumble for their shoes and coats, grab their phones (but forget their room key), then try to find the exits through hallways they’ve only seen once, hours ago when they checked in. The elevators are disabled, so they stumble down 14 flights of stairs with a crowd of other half-awake, disgruntled guests. They gather outside and wait for someone to come and tell them whether it’s safe to go back in.

For the guests, this is a disruption at best and a crisis at worst. Their night has been upended by something they didn’t expect and can’t control, and they’re wondering how much sleep they’ll get before their big meeting tomorrow. All they can do is stand outside in the cold, and hope someone else fixes it soon.

The hotel staff do what they can: directing guests toward exits, calling 911, meeting the fire department at the entrance. But they’re a handful of people with limited training managing a building full of guests who are confused, annoyed, and frightened.

The guests and staff react.

Now think about what happens when the fire department arrives.

The firefighters respond.

Before the first crew even steps off the fire engine, their officer gets on the radio: “Engine 4 on scene, nothing showing, investigating.” Then they check the alarm panel, do a size-up, and start working through well-practiced procedures they’ve followed so many times that they’re second nature. More units are on their way, and even more are just a radio call away, if needed. For the fire department, this isn’t a crisis. It’s just another call in a routine shift.

The difference between reacting and responding isn’t about who cares more. The fire department cares deeply about the safety of the people in that building. The difference is preparation. The guests have no plan, no training, no tools for this situation; they can only react. The fire department has all of those things; they can respond.

The same pattern shows up in software incidents

When something breaks at 2 AM and your on-call engineer gets paged, what happens next? Do they poke at the problem alone, hoping they can fix it before anyone notices? Do they post a vague message in Slack, and then three different people start digging into the same thing without coordinating with each other? Does a senior leader show up and start barking orders, whether or not they have context?

That’s reacting; it’s what happens when people encounter a problem they haven’t prepared for. It’s the natural result of not having a plan, roles, and practiced procedures in place.

Responding looks different. Someone pages an incident commander (IC), and an incident gets declared. The IC assesses the situation and sets initial priorities. Responders are assigned to specific tasks. Communication flows through well-understood channels. Status updates go out at regular intervals. People know what their role is, what’s expected of them, and how to work together effectively in an emergency.

Most engineering teams are full of smart, committed people. What separates chaos from a coordinated response is preparation.

A diagnostic question for your organization

This distinction is one of the most useful questions you can ask about your organization’s incident management maturity: when something goes wrong, does your team react, or respond?

Here are some signs you’re still reacting:

There’s no clear moment when “normal work” shifts to “incident response.” People gradually realize something is wrong and start working on it individually, without explicit coordination.

There’s confusion about who’s in charge.

Multiple people investigate the same thing without knowing it.

Status updates happen sporadically, if at all.

Senior leaders don’t know what’s happening and start asking questions that pull responders away from the work.

When it’s over, nobody is quite sure whether or when to stand down.

Responding, by contrast, has clear transitions: a declaration that shifts the team into a different operating mode, defined roles that people step into, communication practices that keep everyone informed, and an explicit close-out that tells people the emergency is over and they can return to their regular work.

The shift from reacting to responding is incremental, not instant

Most organizations start out reacting. That’s natural. You can’t respond to something you haven’t prepared for, and most organizations don’t invest in incident management preparation until they’ve been burned by a few incidents that didn’t go so well.

The good news is that you don’t need to build all of this overnight. Start with the basics: a clear way to declare that an incident is happening, someone designated as the IC, and a shared communication channel for the incident. That alone will move you from pure reaction toward coordinated response. Then build from there, adding structure, process, and tooling as your team gets comfortable with each new piece.

Your first few formally managed incidents will feel awkward and clunky. That’s fine. The fire department’s recruits feel that way on their first calls too. What matters is that you’re building the capability, one incident at a time. Every incident you manage with even a basic structure is a repetition that makes the next one smoother.

The goal isn’t perfection. It’s preparation. Because when the fire alarm goes off (and it will), the question is whether your team is prepared to respond instead of react.

I’m writing a book about building these capabilities: Incident Management for DevOps and SRE, a practitioner’s guide to structured, effective incident response. If you’d like to hear when it’s available, sign up at im4ds.com.

And if your organization needs immediate help building these capabilities, well, that’s what my consulting practice at Great Circle is all about.

The social contract of emergency mode

When a fire engine rolls down the street without its lights and sirens on, it’s just a big red truck with a fancy paint job. It follows the same traffic rules as any other vehicle its size. Almost nobody pays it any special attention.

But the moment the crew gets dispatched to a call and flips on the lights and sirens, the rules change, not just for the firefighters, but for everyone around them. Other drivers pull over, and cross traffic yields at intersections (supposedly, anyway). Everyone understands that a different set of rules is now in effect, and that those rules are temporary. When the lights and sirens shut off, normal rules resume.

Software organizations need the same kind of shift when an incident happens, but we have to create the signal ourselves, because we don’t have lights and sirens (most of us, anyway; ask me sometime about driving Code 3 and going double the speed limit… at Burning Man, where the speed limit is 5 miles per hour). That signal is the explicit declaration of an incident.

Two modes, one organization

Most technology organizations operate day-to-day in what I think of as “normal mode.” Decisions are made through discussion, deliberation, and consensus-building. Organizational structure follows reporting relationships and seniority. Time is measured in weeks, months, and quarters. This is the right way to run a software organization most of the time.

But when customers are being impacted by an outage, normal mode doesn’t work. You can’t spend three days building consensus on whether to roll back a bad deployment while your service is down. You can’t route a decision through two layers of management while errors are piling up.

When you declare an incident, you shift into “emergency mode.” Time gets measured in minutes and hours. A temporary organizational structure takes effect, where an incident commander (IC) is in charge regardless of where anyone sits in the everyday org chart. A mid-level engineer serving as IC might be coordinating the work of senior engineers and directors. That’s not a problem; it’s the design working as intended. And decision-making becomes more directive: the incident commander makes decisions after considering input, but doesn’t wait for perfect consensus.

This shift feels uncomfortable if you work in a culture that values flat structures and consensus-building. Good. It should feel uncomfortable as an everyday way of operating. Emergency mode isn’t a better way to run an organization. It’s less inclusive, less thoughtful, and more prone to blind spots. But decades of experience in public safety and other fields that deal with emergencies have shown it’s the most effective way to get through an emergency, so you can return to your normal, more collaborative way of working as quickly as possible.

Turning the lights and sirens on (and off)

Here’s the part many organizations get wrong: they never make the shift explicit.

Without a clear declaration that an incident is underway, you get mismatched expectations. Some people treat the situation as an emergency while others respond at their leisure. Some people feel intense urgency while others don’t realize there’s a problem at all. The response becomes uncoordinated, not because people don’t care, but because they’re operating under different assumptions.

Declaring an incident is how you turn on the lights and sirens. It tells everyone: different rules are now in effect. Expect faster communication. Expect a temporary organizational structure. Expect more directive decision-making. This is not how we normally operate, and that’s intentional.

But declaring the end of an incident is just as important as declaring the start. Without an explicit “all clear,” responses fizzle out instead of ending cleanly. People aren’t sure whether they’re still expected to be available at incident-level speed. The on-call engineer who was paged at 2 AM doesn’t know whether they can actually go back to sleep. The incident channel stays open for days with low-priority chatter that nobody feels empowered to shut down.

The IC should explicitly close out the incident: acknowledge everyone’s contributions, confirm the service is restored, and point people toward the post-incident review. This gives responders permission to disengage and return to their normal work. It creates a clear boundary between emergency and normal mode. And it preserves emergency mode as something meaningful, not just a more chaotic version of everyday operations.

The contract

This shift between modes is, at its core, a social contract. You’re asking people to operate differently for a while: to make faster decisions with less information, to set aside some of their usual ways of working, to accept more directive leadership. In return, you’re promising that this is temporary, that it’s only happening because it’s genuinely necessary, and that you’ll return to normal as soon as you can.

The contract breaks down in predictable ways. If leadership declares incidents for things that aren’t really emergencies, people stop taking incident declarations seriously. If responders refuse to shift their behavior during actual emergencies, insisting on the same level of debate they’d use for a design review, incidents drag on unnecessarily. If ICs don’t declare when incidents are over, people get stuck in emergency mode and burn out.

And there’s a particularly damaging failure mode: organizations that are always in emergency mode. If everything is an emergency, nothing is. People become numb to the urgency. They stop responding with the focus and intensity that real emergencies demand. The social contract erodes, and when a genuine crisis hits, the organization discovers it’s lost the ability to shift gears.

Making it work

The most capable organizations I’ve worked with treat this mode-shifting as a core discipline. They declare incidents explicitly. They staff defined incident roles, like incident commander, from a trained pool of people who have other jobs most of the time. They operate under emergency rules for as long as needed, and not a minute longer. And they return to normal deliberately, not by just letting things wind down.

This isn’t about having perfect processes or expensive tooling. It’s about your organization agreeing, in advance, on what emergency mode looks like, when to invoke it, and how to exit it. It’s about building the muscle memory to shift gears when customers need you to, and the discipline to shift back when the emergency is over.

The ability to make this shift cleanly and confidently is one of the clearest markers of incident management maturity I’ve seen across the organizations I’ve worked with. Getting it right changes how your team experiences incidents: from chaotic and draining, to intense but manageable.

This is one of the ideas I’m developing in my forthcoming book, Incident Management for DevOps and SRE. Sign up at im4ds.com to be notified when the book is available, and to receive occasional progress updates and early access to selected content. If your organization needs help with incident management right now, my consulting practice is GreatCircle.com/im.