Spontaneous swarming of responders might seem like a nuisance that breaks our tidy mental models of incident response, but it’s actually very powerful. It’s something to facilitate and encourage, not simply tolerate.
Most of us expect incident response to be fairly linear: an on-call engineer gets paged, decides an incident is called for, and pages an incident commander. The IC spins up an incident channel and pages a couple of additional responders. So far, so good.
What surprises many people is that, within minutes, a bunch of additional reinforcements often begin swarming the incident, without having been paged. A leader at one company I worked with described this as their “white blood cell response”: something goes wrong, and people spontaneously converge to help however they can.
Why do these volunteers show up? For different reasons. Some are engineers who recognize the symptoms from something they’ve seen before, or who know that their systems might be involved. Some are from customer care, account management, or product, keeping an eye on the incident to see if it’s going to affect their areas of concern.
Many of them aren’t there to actively investigate; they’re monitoring, ready to speak up if they have something to contribute. This is one of the advantages of running incidents in Slack rather than on a Zoom call: people can follow along and speak up without disrupting the response. They don’t have to wait for a pause in conversation or work up the courage to interrupt, which in my experience causes many useful observations to go unsaid.
What these spontaneous responders all have in common is that they’re busy experts who have decided that the incident is worth their attention. That’s an organizational strength worth recognizing, developing, and leveraging.
So how do you facilitate and encourage swarming? Start by understanding how incidents actually unfold.
The model vs. the reality
Most incident management processes are built around a tidy model: something happens, an incident is declared, the incident commander pages the right people, those people respond, and the organized response proceeds from there. It’s a useful model, and it describes the formal structure of incident response well enough, but like all models, it has its limits.
Real incidents are messier than that. People find out about problems in all kinds of ways. Engineers start investigating before an incident is even declared. Volunteers show up before the IC has finished their first round of pages. By the time the formal response is organized, a significant chunk of the actual response is already underway, whether it’s coordinated or not.
Often the swarm forms without any incident tooling involved at all. An engineer investigating a problem asks a coworker for help. That coworker pulls in another. Before long, there’s an ad hoc team working on an urgent, significant problem in whatever random place the discussion started, usually some team’s day-to-day channel, where it’s intermingled with “where are we going to lunch today?” and bot posts about new code PRs and deploys. Eventually someone asks “hey, maybe we should be treating this as an incident?” (Pro tip: the answer to that question is almost always an emphatic “yes!”)
When I was leading incident management at Slack, we maintained a standing channel called #incident-next, ready for whoever needed it. Anyone investigating a concern could direct people there to start collaborating in a focused space, separate from whatever team channel the discussion started in. When an incident was declared, the channel got renamed to become the incident channel, preserving all the context, the investigation history, and the people who were already there, and a new empty #incident-next was created to be ready for the next one. This is a simple mechanism that any company could implement today, and it directly supports swarming: people converge, start investigating, and the formal incident declaration catches up to what’s already happening. The response is already underway when the structure arrives to support and harness it.
Most companies design their process only for the tidy model and then get frustrated when reality doesn’t match. A better approach is to design for how incidents actually unfold, which means designing with swarming in mind as a possibility, or even something to be encouraged.
Why swarming works
In a purely command-driven response model, someone has to figure out which teams are relevant, page the right on-call engineers, wait for them to respond, brief them on the situation, and assign them work. Whether that’s the incident commander or a tech lead, every step is bottlenecked through one person’s knowledge and judgment. If they page the networking team but the problem turns out to be in the caching layer, time is lost while they regroup and page a different team. They can only route expertise as well as they can diagnose the problem, and in the early minutes of an incident, the diagnosis is often wrong.
Swarming bypasses that bottleneck. The people with relevant knowledge self-select in, often before the IC or TL even knows their expertise is needed. The engineer who made a related change earlier that day, the one who saw a similar failure pattern last month, the one who happens to know that the caching layer was acting strangely yesterday, they show up because they can see that something is wrong and they recognize that they have something to contribute. Nobody leading the response could have paged all the right people that quickly, because nobody could have known who the right people were.
On a small incident, this is often all you need. A handful of knowledgeable people converge, naturally divide the work, and the problem gets solved quickly, sometimes faster than a formal response could even get organized.
On larger incidents, swarming gets the right people into the room faster than any paging process could. The IC still coordinates the response, but instead of spending the first forty minutes figuring out who to call and waiting for them to respond, they’re organizing people who are already there and ready to work. That’s a fundamentally better starting position. The trick is to make sure that the room is ready for the swarm.
When it isn’t, the same instinct that drives swarming produces freelancing instead: people investigating on their own, without coordinating, because there was no organized response to join. Same motivated people, very different outcome.
Facilitating the swarm
If swarming is this effective, why doesn’t it work well at every company? First, people have to feel safe showing up uninvited. In a culture where volunteering in an incident means risking blame if things go sideways, the swarming instinct gets suppressed before it starts. That’s a deeper problem than process design can solve, but it’s worth naming.
Assuming the culture supports it, two practical things determine whether swarming works:
Can people find the response? When someone hears about a problem, whether it’s a teammate mentioning that checkout is broken, an alert in their team’s channel, an error spike on a dashboard, do they also hear that there’s an organized response underway?
Maybe this is an #all-incidents channel where every declared incident is automatically announced with a link to the incident channel. Active incidents surfaced on the dashboards people are already checking. A predictable naming convention for incident channels, so that anyone can type #inc- in Slack and use auto-complete to see what’s active right now. Integration with your chat platform so that people who are talking about the problem in their team channels get a pointer to the organized response.
The goal is to make sure that everyone who hears about the problem, however they hear about it, also learns where people are converging.
Then, can they easily contribute? Once people find the channel where the response is happening, they need a clear on-ramp. A pinned message with the current situation, who’s filling what role, and what’s being investigated. A clear expectation that new arrivals announce themselves and what expertise they bring, then self-brief by reading the channel and the latest situation report before asking for an assignment. Enough structure that someone can get oriented in a few minutes and start contributing, rather than facing a wall of scrollback with no obvious way to plug in.
None of this is complicated. But it does require deliberate design. The swarming instinct happens naturally. The infrastructure that makes swarming productive doesn’t.
Invest in the instinct
If your team is already swarming incidents, you’ve got something invaluable: people who care enough to show up. The rest is infrastructure: making the response visible, making it easy to join, and giving the IC the tools to organize a growing swarm on the fly.
You can’t manufacture that instinct. But you can make it dramatically more effective.
Swarming is one of the topics I cover in depth in my forthcoming book, Incident Management for DevOps and SRE. If you’d like to know when it’s available, sign up at im4ds.com. And if your company needs help with incident management right now, my consulting practice is greatcircle.com/im.
I used to do something similar at work. We used IRC rather than slack. We had tac1, tac2, tac3, and tac4 IRC channels with supybot already deployed.
They were used as a low barrier method of coordinating an incident or scheduled major work.
For example we used the tac channels for our scheduled production migration from one data center to another DC. So we trained the workflow in non-emergency incidents. Similar to how EMS sets up an IC and treatment sectors for large events.
This also gave experience in using supybot to page/invite people from other channels, post particular segments of the IRC chat to other channels or web pages, email parts of the chat, bridge the chat to other channels (for a wider audience).
We usually used TAC 1 but if we had an incident and part of the incident response was to move traffic, tac1 was used to investigate and solve the issue while tac2 was used to coordinate moving the second data center into production state. Tac3 and 4 could be used to troubleshoot other tasks (using supybot to log milestones to other tac channels).