Every fire department has a training program. Big-city departments have entire training divisions; even small volunteer departments that can’t spare anyone full time still name a training officer. Not because training is the department’s mission, but because maintaining the capability to do the mission requires sustained, dedicated attention.

New recruits need to be brought up to speed. Everyone needs to learn about evolving techniques and new equipment. Procedures need to be updated as building codes and materials change. Hard-won lessons from past incidents would survive only as stories told around the kitchen table; the fire service has a strong storytelling tradition, and its legends and cautionary tales carry real value, but oral history is hard to study, standardize, and train on.

Maintaining operational capability is itself a job, distinct from the operational work it supports, and fire departments size the role to the department rather than leave it unassigned.

Many software companies haven’t learned this yet. They invest real effort in building an incident management process. They define severity levels, write runbooks, designate incident commanders (ICs), set up communication channels. The project might take weeks or months of focused work, often driven by someone who cares deeply about doing it right (and often done in their “spare time”). When it’s done, it works, at least for a while. Incidents get declared. ICs run the response. Post-incident reviews happen. Everyone takes it for granted.

Then the person driving it gets promoted, or moves to another team, or leaves the company. The process, which was never really institutionalized because it didn’t need to be while that person was carrying it, begins to decay. Not catastrophically, but more like a garden nobody is tending any more: it doesn’t collapse overnight, it just slowly fills with weeds until one day you look up and realize the original design is barely recognizable.

The training materials haven’t been updated since the initial rollout. New engineers join but never go through incident training because nobody is scheduling it anymore. The severity level definitions still describe one product, but the company now has three. The IC rotation is running on the same six people it started with, even though the engineering team has doubled in size. The post-incident review template still references a tool the company stopped using a year ago.

None of these are crises on their own. Each one is easy to defer. But they compound, and the cumulative effect is that the process on paper bears less and less resemblance to what actually happens during incidents. In my experience, six months is roughly how long institutional momentum carries before the absence of active stewardship becomes visible in the quality of your incident responses. And growth accelerates the decay: the company simply grows away from the process, and nobody’s job is to notice.

This is what happens when you have a process but not a program.

A process is not a program

A process is a set of documented procedures: how incidents get declared, who fills which roles, what communication channels to use, how to run a post-incident review. A process can be written down, trained once, and followed.

A program is the organizational structure that develops, maintains, evolves, and champions the process over time. It’s the thing that keeps the process alive.

Many companies build the process and assume they’ve built the program. They haven’t. They’ve written a document, and documents don’t train new hires, don’t recruit for on-call rotations, and don’t update themselves when the company reorganizes around them. People do those things, and it only happens reliably when it’s actually somebody’s job.

“Everybody owns it” means nobody owns it

When I ask companies who owns their incident management program, the most common answer is some version of “we all do” or “the engineering organization as a whole.” This sounds collaborative. In practice, it means nobody has the explicit responsibility, the dedicated time, or the institutional authority to keep the process alive.

This organizational challenge isn’t unique to incident management. Companies that are serious about security don’t say “everybody owns security” and leave it at that. They assign ownership because shared responsibility without explicit ownership means the work doesn’t get done.

Incident management is the same kind of organizational capability. It needs someone whose actual job, not just their passionate side interest, is keeping it healthy.

What a program actually does

When I talk about an incident management program, I mean ownership of the full lifecycle of the capability, not just the procedures themselves. That includes keeping everything current as the company grows and changes: process documentation, severity definitions, escalation paths, tooling, runbooks.

It includes running a training pipeline so new hires are prepared before their first real incident, not thrown into the deep end during it. It includes maintaining the incident commander corps: recruiting new incident commanders, nurturing their development, supporting healthy on-call rotations across teams, and recognizing the people who do this demanding work. My former Slack colleague Scott Nelson Windels likens this to the farm teams and academies that elite sports clubs run: the point isn’t just fielding today’s roster, it’s making sure capable players are always coming up to fill it next quarter, too.

And it includes owning the post-incident review process and looking across incidents for patterns that no individual team would spot on their own. It includes tracking whether the process is actually being followed, and investigating when it isn’t, not to punish people, but to understand whether the process needs to change.

No single component is enough on its own, and no component stays healthy without sustained attention.

The good news

Building a program doesn’t require hiring a large team or creating a new department. At many companies, especially smaller ones, it starts with one person who has explicit ownership and dedicated time. What matters is that the responsibility is named, visible, and institutionally supported, not just assumed.

Here’s a quick test. Ask who owns your incident management program. Not who wrote the process, and not who ran the last big incident, but who is accountable, today, for whether the training is current, the rotations are staffed, and the severity levels still match the product. If the answer is a name, the follow-up question is what happens when that person leaves. If the answer is “everybody,” or someone who left the company last year, the process is quietly withering. And if you have a program but it would collapse without you, you haven’t finished building it yet.

The fire department didn’t name a training officer because it had extra budget. It named a training officer because it understood that maintaining a capability requires ongoing investment. The alternative, assuming trained firefighters stay trained and procedures stay current without anyone specifically owning those things, is how capabilities quietly erode until they fail when you need them most.


I’m writing a book on Incident Management for DevOps and SRE. If you’d like to know when it’s available, and get occasional updates along the way, you can sign up at im4ds.com.

If your company needs help building its incident management program, that’s the focus of my consulting practice at Great Circle.