How senior leaders disrupt incidents just by showing up

A vice president joins an incident channel and posts a single, reasonable question: “What’s the customer impact?” Every responder in the channel pauses. The person investigating the network layer stops to check whether they should be the one to answer. The engineer who was about to try a promising mitigation hesitates, wondering if the VP’s question implies a different priority. The incident commander (IC) now has to decide whether to answer the VP or redirect them, and either way, the response has lost momentum.

There’s nothing wrong with the question itself. If a fellow engineer had asked it, nobody would have blinked. The problem is who asked it and where. A VP’s question in the main incident channel doesn’t land the way a peer’s question does. It lands with the weight of the org chart behind it, and everyone in the channel feels it. If the VP had sent the same question privately to the IC, or to an executive liaison, it would have been answered without disrupting anyone. Instead, it went to the room.

This is a familiar problem in incident management, and VPs are the usual example, but they’re not the only ones who cause it. Anyone who carries organizational weight can have the same effect: directors, senior architects, the principal engineer who designed the system that’s currently on fire. The dynamics are the same; only the job title changes.

The aircraft carrier in the harbor

A useful way to think about this is to picture an aircraft carrier entering a harbor. The carrier isn’t doing anything wrong. It’s not speeding or behaving recklessly. But it’s enormous, and everything else in the harbor has to adjust: smaller vessels change course, dock operations pause, harbor traffic rearranges itself around the carrier’s presence. The disruption isn’t caused by anything the carrier does. It’s caused by what the carrier is.

Senior leaders have the same effect in incident channels. When a VP joins and posts a message, responders notice. People stop what they’re doing to read it. Some start formulating answers, even if the question wasn’t directed at them. Others worry about what the VP’s presence means: Is the response not going well enough? Are we in trouble? Should I be doing something different?

None of this is the VP’s intent. They just wanted to understand what was happening. But the effect is real, and it’s disruptive in ways that senior leaders often don’t recognize, because they can’t see the disruption they’re causing from where they sit.

The disruption you can’t see

This is what separates the presence problem from the more obvious forms of executive disruption. The conventional advice focuses on the visible behaviors: a VP overriding the IC’s decisions, an executive asking rapid-fire questions that pull responders off their tasks, someone senior giving orders that conflict with the tech lead’s plan. Those are real problems, and they deserve attention. But they’re fixable with straightforward norms: address questions to the IC privately, don’t give orders in the main channel, defer visibly to the incident leadership.

The presence problem is harder, because it persists even when the senior leader does everything right. A VP who joins the channel and says nothing still changes the room. Their presence will be noticed, even if Slack doesn’t helpfully snitch on them with a “so-and-so has joined the channel” message to the channel. The presence of someone senior often introduces a layer of self-consciousness that slows things down, even when that person’s intent is purely to observe.

The really tricky cases are the senior leaders who are also genuine technical contributors. At one company, I worked with two senior executives who’d been there since its earliest days. Both were talented engineers who often made real technical contributions during incidents. They weren’t barging in to ask uninformed questions; they were explaining old code and spelunking through logs and spotting things that less experienced responders would have missed.

Their contributions as subject matter experts were unquestionably valuable. But every time they showed up in an incident channel, many responders (especially newer hires who hadn’t worked side-by-side with them for years) reacted to their titles, not their expertise. The disruption was so tied to their identity that I remember half-seriously considering whether I should tell them to set their Slack display names to secret identities (“Hal Jordan” and “Diana Prince”?), so they could contribute as “just engineers” without anyone knowing The Boss was in the room.

We never actually did it, but the fact that “give them secret identities” was the best solution anyone could think of tells you something about the nature of the problem. It wasn’t what they were doing; it was who they were.

And it’s not limited to people with management titles. When I was at Slack, I was an individual contributor with no direct reports, but I led the incident management program and had trained most of the responders and nearly all of the incident commanders. If I joined an incident channel and started asking questions or making suggestions, some would react to my presence the same way they would to a wandering VP. I had to be very intentional about how I showed up, and I probably wasn’t always as careful as I should have been.

The test isn’t your title; it’s whether your presence changes the room.

What senior leaders can do

The solution isn’t to exclude senior leaders from incidents entirely. Some, like the executives in the story above, are genuine technical contributors whose expertise makes the response better. Others have legitimate information needs: they may need to brief the board, reassure a key customer, or make business decisions that depend on when service will be restored. Both cases deserve to be handled well, but they need different approaches.

The first question to ask yourself is, do you really need to be visible there at all? Your expertise may be valuable, but so is a response that isn’t reacting to your mere presence. If the answer is yes, be deliberate about how you enter: explicitly state your role (“I’m here as an SME on the payments system; Alex is still the IC”), remind people to take their direction from the incident leadership rather than from you, and consider setting your display name to reinforce it (both Slack and Zoom let you do this). And once you’re in the channel, be disciplined about staying in your stated role. An executive who joins as a subject matter expert but starts asking strategic questions has reintroduced the problem through the back door.

On the other hand, if you just need to stay informed and make business decisions, your goal should be to do that in a way that doesn’t visibly put you in the incident channel.

Get your information from the sitreps. A well-run incident produces periodic situation reports (sitreps) that are designed to answer exactly the questions senior leaders have: what’s happening, what’s the impact, what’s the plan, and when’s the next update. If the sitreps aren’t meeting your needs, that’s something to work with the team on between incidents, not by visibly disrupting the current incident.

If you must communicate with the response, go to the IC privately. Send a private message. Don’t post in the main channel, even to “just ask a quick question.” There’s no such thing as a quick question from a senior leader during an incident.

Consider establishing an executive liaison role. On larger incidents, one senior leader can serve as the conduit between the response and the rest of the leadership team. The liaison works directly with the IC, passing information along to the other leaders and representing their concerns, so that nobody else on the leadership team needs to enter the incident channel. This is one of the most effective structural solutions to the presence problem.

Run interference, in coordination with the IC. One of the highest-value things a senior leader can do during a major incident is handle the organizational demands that would otherwise land on the IC: the sales team asking what to tell a customer, the legal team needing clarification, the product team wondering whether to delay a launch. But this only works when it’s coordinated with the IC, not freelanced. A quick conversation (“I’m going to handle incoming questions from sales and legal so they don’t land on you; I’ll use the latest sitrep as my source”) keeps the IC informed and avoids the risk of a senior leader making commitments that contradict what the response team is communicating.

What ICs can do

Even in companies with good norms, senior leaders will sometimes show up in the incident channel. The IC needs to be prepared for that.

When a director drops a question into the channel, intercept it before responders start trying to answer. A simple “Thanks, I’ll follow up with you on that directly” takes the question out of the channel and signals to responders that they should stay focused on their assigned tasks. This is the IC acting as a shield, absorbing the disruption so the responders don’t have to.

Proactive communication also helps. If you write clear sitreps that include a timestamp and an expected time for the next update, readers can judge how fresh or stale the information is without having to ask. And if you consistently meet the update cadence you’ve committed to, that builds confidence over time that the updates will keep coming.

When senior leaders can trust that the IC will keep them informed as the response progresses, they’re less likely to come hunting for information themselves. It takes time, over several incidents, to build that trust, but it’s one of the most valuable investments an IC can make.

Awareness is the first step

The hardest part of this problem is that the person causing the disruption almost never sees it. From the bridge of the aircraft carrier, the harbor looks fine. It’s the smaller vessels that changed course. That’s why this isn’t just a behavior problem to be solved with a list of don’ts. It’s an awareness problem. The companies that handle this well aren’t the ones with the most rules about executive behavior during incidents. They’re the ones where senior leaders have internalized a simple idea: during an incident, the most helpful thing you can do might be to stay out of the way, while the most disruptive thing you can do is show up with the best of intentions.

When declaring an incident becomes everyone’s favorite workaround

You see someone declare a Sev-2 and you wonder: wait, why is that even an incident? Nothing is down. Customers aren’t affected. But a manager needed to get their team’s problem to the top of another team’s priority queue, and the incident process was a reliable way to make it happen. That’s not really what the incident process is for, but it worked, so where’s the harm?

The problem is, once folks see that this works, it starts happening more often. A product manager declares an incident because the incident notification is the fastest way to get leadership attention on a problem that’s been stuck in the backlog for weeks. An account team declares one because they need engineering support for a big demo to a major prospect and the incident process is the easiest way to pull engineers out of their sprint work on short notice. An engineer declares one because it’s easier than navigating the formal exception process for the deployment freeze.

The harm is cumulative. When a growing fraction of your declared “incidents” aren’t real emergencies, the urgency signal degrades. When a genuine Sev-1 arrives, people respond with less urgency because they’ve been conditioned to expect another workaround. And the incentives compound: folks who game the system get their problems solved faster, which teaches everyone else that gaming is how to get things done. Each individual declaration is an understandable decision by someone who needs to get something done; it’s the aggregate that corrodes the process.

Every one of these non-emergency declarations still carries the full overhead of a real incident. Responders get pulled off their planned work. Someone drops whatever else they were doing to serve as incident commander. Stakeholders context-switch to follow along. When you’re running enough of these, your teams are spending a meaningful fraction of their time in emergency mode for things that aren’t really emergencies, and all the indirect costs of incidents (disrupted projects, context-switching, recovery time) accumulate just the same.

There’s an irony here: people are reaching for the incident process because it works; they’ve seen that it reliably delivers coordination, prioritization, and urgency on demand.

The instinctive response is wrong

When companies notice this pattern, the instinctive response is often to tighten the declaration criteria. They add gatekeeping: maybe you need manager approval to declare an incident, or there’s a pre-declaration checklist you have to complete first, or someone reviews whether the declaration was “warranted” after the fact. The intent is reasonable. The net effect is corrosive.

Gatekeeping incident declarations is counterproductive. Every speedbump you build also slows down real incidents. The person who hesitates to declare because they’re not sure the problem is “bad enough” is already a common failure mode in incident response. Adding a formal approval step or a post-hoc review of whether the declaration was justified makes that hesitation worse, not better.

You also miss what the gaming is telling you: people reaching for the incident process are telling you that your normal processes are falling short. If you only crack down on the gaming, you suppress the symptom without learning anything from it, and the underlying problems persist.

Fix the escape routes, not the escaping

Instead, look at what side effects people are trying to trigger when they declare questionable incidents, and make those capabilities available through other means.

If the easiest way to bypass the deployment freeze is to declare an incident, create a non-incident exception process for urgent changes. This doesn’t have to be complicated; a lightweight approval from a designated release manager, with a clear escalation path, covers most cases.

If the easiest way to get your problem moved up another team’s priority queue is to declare an incident, create a prioritization escalation path that doesn’t require an incident. A cross-team triage meeting, an explicit expedite-request mechanism, or even a dedicated Slack channel that the right people actually monitor can absorb most of the pressure. The bar doesn’t have to be as high as “declare an emergency”; it just has to be lower than “wait six weeks for the next planning cycle.”

If the easiest way to assemble a cross-functional team on short notice is through the incident process, create a lightweight coordination mechanism for non-incident situations. Some companies call these “swarms” or “tiger teams” or “coordination requests.” The name doesn’t matter; what matters is that people have a way to get the collaboration they need without borrowing the incident process to do it.

Repeatedly gaming the incident process to get resource prioritization or cross-functional coordination isn’t a series of one-off workarounds; it’s a symptom of a systemic problem that needs a systemic response. Google’s SRE organization built formal Code Yellow and Code Red mechanisms for exactly this: structured ways to rally resources and elevate priority when a problem is serious enough to demand cross-functional attention, but isn’t an incident.

The diagnostic question

Look at your last dozen or so incidents and ask, for each one: was this declared because there was an emergency, or because the incident process was the easier path to something the team needed?

You don’t need a formal audit. Just ask a few experienced incident commanders and on-call engineers; they already know which ones were real and which ones weren’t. Then talk to the folks who called for the questionable ones (in a blameless, fact-finding way, of course). They’ll tell you exactly what’s missing from the normal processes, if you’re willing to listen.

People gaming the incident process is just a symptom. The underlying problem is usually that normal processes are too rigid, too slow, or too unresponsive, and the incident process is the path of least resistance. Fix the underlying problem and the gaming stops, because there’s nothing left to game around. Your incident urgency signal recovers, your teams stop burning emergency-mode cycles on non-emergencies, and when a real Sev-1 hits, people respond like it matters.

And if you’re dealing with this, take a moment to appreciate what it says about your incident process: people are borrowing it because it works. The fix isn’t to make it stop working. It’s to make everything else work that well too.

Incidents start before the response does

Your company has probably invested significantly in what happens after an incident is identified: incident response tooling, trained incident commanders, communication protocols, on-call rotations. That investment matters. But what about the gap between when a problem starts and when anyone on your team knows about it?

During that gap, customer damage is accumulating. The problem is getting worse, the blast radius is expanding, and nobody on the team is doing anything about it because nobody knows yet.

You can’t eliminate this gap entirely, but you can shrink it. Four investments make the biggest difference.

Broaden your detection surface

Automated monitoring is the first and best line of defense, but it can only catch the failure modes someone thought to check for. Human detection isn’t a gap you can eliminate; it’s a permanent and valuable part of your detection capability.

This means your customer support team is part of your detection infrastructure, whether or not you’ve told them so. So is any part of your company that interacts with customers regularly: account execs, customer success managers, even your social media team. They talk to your customers every day and often see concerns emerge before engineering does. And don’t overlook your customers themselves, who won’t limit their reports to your “official” support channels. If all these folks don’t have clear, fast escalation paths to flag potential problems for engineering, you have a detection gap that no amount of monitoring investment will close.

If your company is a heavy user of its own product, the detection surface extends even further. When I led Slack’s incident management program, literally anyone in the company might notice a problem while using Slack internally. Not every company is in that position (it depends entirely on what the product is), but those who are should take advantage of it. Make sure everyone (all the way down to the part-time security guard covering the front desk on weekends) knows how to report problems they see.

And watch for indirect signals. One of Slack’s best harbingers of “something is broken, even if we don’t know what yet” was the page-view rate on our public status page. If it started surging upward, we knew that something was wrong, even if we weren’t getting any other clear signals yet, and we’d start investigating. It was like smelling a light waft of smoke, well before the smoke detectors and fire alarms go off. If you have a public status page, consider adding its traffic patterns to your monitoring. A sudden spike in visits is a low-cost early warning powered by the collective behavior of your user base.

Lower barriers to reporting

Most of these detection channels depend on someone raising a concern, and that only works if the barrier to doing so is low. At many companies, the only mechanism for raising an alarm is to declare an incident, which triggers a full coordinated response: pages go out, a channel is created, an incident commander is assigned, people drop what they’re doing.

That’s appropriate when you know you have a real problem. But if the only way to raise a concern is to trigger that entire response, people will hesitate, and rightfully so. Nobody wants to be the person who launched a full incident response over a hunch that turns out to be wrong. So they wait for more evidence, and the detection gap grows.

Think of it like calling emergency services. When you call 911 (or 999, 000, 112, or whatever your country’s emergency number is), you don’t have to know whether you need an ambulance, a fire engine, a hazmat team, or a bomb squad. You describe what you see, and a trained dispatcher determines how serious the situation is, what sort of response is warranted, and who to send.

Your incident detection should work the same way: make it easy for anyone to say “I think something might be wrong,” and let someone with training, experience, and context determine what response is warranted. At Slack, introducing a lightweight mechanism for exactly this was one of the most impactful things we did.

Continuously right-size your alerting

It’s tempting to close the detection gap by making your monitoring more aggressive: lower the thresholds, add more alerts, page on anything that twitches. This can backfire badly. Every alert that wakes someone at 3 AM and turns out to be nothing makes it a little more tempting for your on-call engineers to dismiss the next one. Alert fatigue is one of the most insidious threats to detection, precisely because it accumulates gradually. Your alerting system doesn’t fail all at once; it erodes, one false alarm at a time, until the real alerts get lost in the noise.

The discipline runs in both directions: yes, add monitoring when you discover gaps, but regularly prune alerts that aren’t earning their keep. If a service-owning team can’t get through a review of every alert they received in the past week in a reasonable portion of a weekly ops review meeting, they’re getting too many alerts.

Examine the gap

Another way to shrink the detection gap over time is to examine it after every incident. You’re never going to be able to fully automate detection, but it’s still an ideal worth pursuing. Three questions, asked consistently in every post-incident review, create a steady stream of improvements:

  • How long was the gap between when the problem started and when we detected it?
  • Could we have detected it sooner?
  • What monitoring would we need to add, or what threshold would we need to adjust, to catch this kind of problem faster next time?

The bottom line

Investing in detection is investing in the foundation of your entire incident management capability. You can have well-trained incident commanders, practiced responders, and polished communication protocols, but none of it matters until you know there’s a problem.