It’s 3 a.m. in California, where most of the dev team are still snug in their beds. The auth system has started rejecting valid credentials. Early bird East Coast customers are already trying (and failing) to log in for the day, and thousands of users in Europe have already given up and gone elsewhere. In a couple of hours, the West Coast will be waking up too. A brilliant engineer swoops in and saves the day. She has legendary debugging skills and a deep understanding of the auth system, and she puts together a fix in forty minutes that would have taken anyone else hours to even diagnose. Later that morning, leadership is sending thank-you messages in the all-hands channel. Her VP awards her a small spot bonus, and her manager reminds her to include it in the next performance review cycle.
What doesn’t usually happen is anyone asking: what if she hadn’t been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.
Near misses look like successes
In aviation and other safety-critical fields, it’s widely accepted that a near miss is an unparalleled opportunity to learn and deserves the same investigation as an actual failure. The reasoning is straightforward: a near miss reveals the same systemic vulnerabilities that a failure does. The only difference between a near miss and a disaster is that the outcome happened to be good this time, often because of luck, timing, or the presence of one specific person.
A heroic incident response is a similar opportunity. The system nearly failed, and would have failed if that one engineer hadn’t been available or hadn’t known exactly what to do. Her skill, expertise, and dedication are worth appreciating. But her unavailability would have meant a much worse outcome, and that’s worth examining too. Too many companies celebrate the save and stop there.
The incentive nobody designed
When a company celebrates a heroic save without examining why the heroics were necessary, it sends a message. The message isn’t intentional, but it’s clear nonetheless: what gets valued is the dramatic rescue, not the boring preparedness work that would have made the rescue unnecessary.
Over time, that message shapes behavior. The engineer who writes thorough runbook documentation, trains new team members on the auth system, and invests in monitoring improvements doesn’t get the same recognition as the one who swoops in at 3 a.m. and saves the day. Preparedness work is largely invisible in performance reviews. Heroic saves are memorable.
The result is a perverse incentive loop. Heroics get rewarded, preparedness doesn’t, and the company remains dependent on heroic saves because nobody is investing in the alternative. This isn’t because anyone explicitly decided that preparedness doesn’t matter. It’s because the reward system is quietly rewarding the wrong thing, and nobody has noticed because the heroes keep delivering results. Until they don’t.
In my experience, this is one of the most common patterns in companies that are struggling with incident management. They have talented, dedicated people who keep delivering heroic results, and because the results keep coming, nobody realizes there’s a growing structural problem underneath.
The hero as single point of failure
The incentive loop creates a second problem. The hero gradually becomes a bottleneck and a single point of failure. When that engineer is on vacation and the next auth system incident hits, the team might spend hours just figuring out what’s going wrong, let alone fixing it. When they eventually leave the company (as they likely will; heroes tend to burn out), the team discovers that critical knowledge walked out the door with them.
I see this pattern regularly in my consulting work. In a company’s most serious incidents, it keeps turning to the same handful of heroic engineers. Those engineers are talented and committed, and their involvement has genuinely saved the company from significant damage. Everyone involved with incidents knows who they are, and breathes a sigh of relief when they join an incident channel. But the company has never seriously examined what its response capability looks like without them. The term that often comes up to describe these people is “indispensable,” which is really another way of saying that the company’s incident response capability depends on specific individuals’ availability.
Why the problem stays hidden
The most insidious aspect of this pattern is that it’s invisible to leadership for as long as the heroes keep delivering. Companies at the earliest stages of incident management maturity often don’t realize they’re at risk. Leadership sees consistently good outcomes and assumes the company has strong incident response, when what they actually have is strong individuals (and a certain amount of good luck).
By the time the fragility surfaces, the gap between where the company thought it was and where it actually was can be startling.
Heroic is a growth stage, not a compliment
When I assess incident management capabilities for my consulting clients, one of the dimensions I evaluate is program maturity: where is this company on the growth path from ad hoc response to reliable organizational capability? The first stage on that path is called “Heroic.” It isn’t meant to be flattering. It means that incident response quality is a property of specific talented individuals rather than a property of the company. When those individuals are available, things go well. When they’re not, things go sideways.
Every company starts here. The question is whether they invest in growing past it, converting individual capability into organizational capability. That transition is what the rest of the maturity model describes, and it’s the core of what effective incident management programs are designed to do.
What to recognize instead
None of this means companies should stop recognizing heroic contributions when they happen. When someone saves the day at 3 a.m., thank them. But also investigate why the heroics were necessary, and invest in the answers. That’s a form of recognition too: it says the save mattered enough to learn from.
To move from “Heroic” to higher levels of organizational capability, you need to shift what gets sustained recognition. Recognize the work that makes heroic saves unnecessary: the runbooks, the training, the well-coordinated responses where nobody had to be heroic.
If an engineer spent much of their quarter writing runbooks, training new responders, and coordinating incident responses, recognize that work: in performance reviews, in public acknowledgment from leadership, in awards and bonuses. If you don’t, you’re telling your organization that the only incident management work worth noticing is the dramatic save.
The goal is to make effective incident response something the company can do reliably, regardless of who happens to be on call. Heroes are still welcome, and still admired. They just shouldn’t be required.
Recent Comments