Ask six people at the same company what counts as an incident, and you’ll likely get six different answers.
The IT ops manager thinks of every ticket in ServiceNow. The on-call SRE thinks of the alert that paged at 2 AM. The head of sales thinks of the call from an angry enterprise customer. The finance team thinks of the disruptions they have to issue SLA credits for. The engineering manager thinks of last week’s outage, the one she was up until 3 AM for, acting as the incident commander and coordinating the response across three teams. The newest engineer isn’t sure, but knows they don’t want to be the one who declares one.
This confusion isn’t academic. When people in the same organization mean different things by “incident,” practical problems follow. The CFO asks “how many incidents did we have last quarter?” to forecast SLA payouts for the board report, and gets a completely different number depending on who answers. Engineering can’t tell whether the trend line is improving because the data mixes SRE pages, ITIL tickets, and multi-team coordinated responses into a single count. Someone declares an incident and half the organization thinks it’s a crisis while the other half thinks it’s Tuesday. And every attempt to improve incident management stalls because the people in the room haven’t realized they’re talking about different things.
Where the Confusion Comes From
The word “incident” has been overloaded with several overlapping meanings in the technology industry; people think they’re all talking about the same thing, but they aren’t, and the confusion causes more problems than most people realize.
ITIL has taught generations of IT professionals that an incident is any unplanned interruption to a service. Under that definition, a single user unable to log in is an incident. A slow database query is an incident. A printer jam is an incident. ITIL-trained teams sometimes process thousands of “incidents” per month through their ticketing systems. When someone from that background hears “we need better incident management,” they’re thinking about ticket queues and resolution targets, not about coordinated emergency response.
PagerDuty, the most widely-used on-call tool in the industry, has long compounded the problem through its product terminology. In PagerDuty’s data model, which stretches back over 15 years to the early days of the company, every alert that pages someone creates an “incident.” You literally cannot page a colleague without creating a PagerDuty “incident.” This means an on-call engineer who gets paged three times on a quiet Tuesday has, in PagerDuty’s language, experienced three incidents. Over time, this trains people to think of “incident” as synonymous with “page” or “alert.” PagerDuty has recognized the gap as they’ve expanded into incident response orchestration, particularly after acquiring Jeli in 2023, adding the concept of “major incidents” (a term ITIL also uses, for the same reason) to distinguish coordinated-response situations from everyday pages. But the underlying product model remains, and generations of engineers have already internalized the equation: incident = someone got paged.
Then there’s the other extreme. Some organizations reserve “incident” exclusively for the worst events they can imagine: full-site outages, data breaches, events that make the news. This sounds disciplined, but it creates a high psychological barrier to declaration. If “incident” means “catastrophe,” nobody wants to be the person who declares one for something that turns out to be just a blip. So people hesitate, they wait for clearer signals before acting, and by the time someone finally says the word “incident,” the situation has been burning unchecked for longer than it needed to.
Those are all legitimate uses of the word, but they’re answering different questions. The ITIL definition tells you what to put in the ticketing system. PagerDuty’s tells you what to attach a page to. Finance’s tells you what catastrophes will trigger an SLA payout. None of them help you decide when to urgently pull people together and mount a coordinated response.
We’re not saying those other definitions are wrong and ours is the One True Definition. We just want folks to be aware of the potential for confusion, and to be sure they understand from context (or explicit clarification) which version of “incident” someone is referring to.
It’s like the word “security,” which means something completely different to an information security engineer, a physical security guard, and a financial analyst trading stocks and bonds. Nobody argues about which meaning is correct; they just make sure everyone in the conversation knows which one they’re talking about.
My professional community, my own work, and my forthcoming book are about managing emergencies that require urgent, multi-person, coordinated responses. In those contexts, we use “incident” to mean a situation that is significant enough to need attention right now, urgent enough that you can’t just file a ticket and walk away, and beyond what one person can handle alone. That’s one valid, context-specific definition among several. The important thing is that everyone in a conversation is clear on which one they’re using in the moment.
Watch for It
Here’s a concrete step you can take this week: ask six people at your company what they think counts as an incident. Don’t ask in the abstract; give them three or four scenarios and ask which ones they’d call an incident. Better yet, do it in a group setting where everyone can hear each other’s answers. The spread will surprise people, and that surprise is the point.
Once your team sees the mismatch, you gain the ability to catch it in the moment. “Wait, are we all talking about the same kind of incident right now?” is a surprisingly useful sentence.
This is one of the communication challenges I explore in my book, Incident Management for DevOps and SRE, which I’m currently writing. If you’d like to be notified when it’s available, sign up at im4ds.com.
If your organization is wrestling with these questions right now and doesn’t want to wait for the book, my consulting practice can help.
Recent Comments