Today’s AWS outage highlights a critical question for all organizations: what’s your backup plan when your regular incident management tools are themselves experiencing disruptions? Slack, Zoom, PagerDuty, incident.io, FireHydrant, Rootly, and Atlassian (Jira, Opsgenie, Jira Service Management) and many others are currently reporting various issues due to the outage.
A big challenge during incidents like this is re-establishing communication among responders. If your usual platforms like Slack and Zoom are down, can you reliably transition to alternatives like Google Meet or Microsoft Teams? The key is to have a pre-defined fallback communication plan that’s known to everyone on the team, not just tucked away in an audit document that nobody remembers exists.
This plan needs to specify the exact meeting ID, channel, or other critical details for reconvening. This information must be accessible even when your primary communication channels and document stores are unavailable. Consider distributing laminated wallet cards with essential phone numbers and web URLs to key staff, or storing this information in a shared password vault that offers offline access. Alternatively, your IT department could ensure this crucial data is kept up-to-date in a file on every employee’s laptop.
Security is important. You have to assume that, over time, this fallback information might become accessible to unauthorized parties (press, competitors, hackers, etc.). Therefore, when using these reassembly mechanisms, avoid discussing highly confidential or sensitive matters in the main fallback channel. Instead, treat the fallback channel just as a rendezvous point, and use it to direct responders to move to specific secure channels where all participants can be carefully identified and authenticated.
These fallback mechanisms need to be exercised periodically, both to ensure that folks remember that they exist and to verify that the specific mechanisms still work as intended.
—
Incidents can be incredibly costly, impacting revenue, sales, reputation, employee productivity, morale, and turnover. Proactive preparation and continuous learning from every incident are vital. As an expert incident management consultant, advisor, and coach, I can help your organization develop these essential skills and prevent costly problems. Contact me today to discuss how I can support your team.
Situation reports (SitReps) are vital for effective incident management. They help align current responders, orient new responders, and keep stakeholders updated. Here’s how to create impactful SitReps:
Consider multiple audiences:
Responders: Provide an overview of the problem and current actions being taken.
Executives/Stakeholders: Provide high-level information on the impact and prospects for resolution.
Make each SitRep self-contained, and keep it short:
A good SitRep should be readable and understandable in under two minutes.
Each SitRep should provide sufficient information for a new team member to understand the incident without needing to read previous reports.
Avoid making SitReps a comprehensive log of everything that has happened in the incident so far; focus on current information, and direct readers to Slack channels or incident logs for a full history.
Remember the mnemonic “CAN” for structuring your SitRep:
Conditions: What is the current situation?
Actions: What are we doing about it?
Needs/Next Steps: What additional resources are required? What are we planning to do next? Sometimes, all you need is time for an in-progress fix to finish rolling out; if that’s the case, just say so.
Maintain a predictable cadence:
Establish a regular schedule for SitReps (e.g., hourly for availability incidents) so readers know when to expect updates and how current the information is.
Include a timestamp on each SitRep, so readers can judge how “fresh” it is.
Explicitly state when the next SitRep is expected to be shared. Provide a specific timeframe, such as “next update in 60 minutes”; avoid vague statements such as “next update when circumstances change.”
Publish early for significant changes, but never delay a scheduled SitRep, even if there are no major updates.
Leverage Slack features and facilitate sharing:
If you’re using Slack for incident communications (as you should be; see my previous blog post Why Slack outshines Zoom for incident management), pin the most recent SitRep in the incident channel.
Train responders and observers to check pinned posts when they first join an incident channel.
Unpin older SitReps when new ones are pinned to keep the number of pinned posts manageable.
Create each SitRep as a separate message or document to facilitate easy forwarding, sharing, and pinning.
Guide follow-ups:
Clearly state who to contact and how for questions or corrections (e.g., DM the Incident Commander, post in the incident channel).
Standardize with templates:
Develop a consistent SitRep template for use by all incident commanders and communication leads in order to save time, ensure comprehensive coverage, and provide a familiar format for readers.
Here’s an example of a good SitRep:
SitRep 07-Jun-2025 13:35 PDT (20:35 GMT) #i-5150-api-timeouts Active Sev-2 IC @brent
Situation: Since about 10:30 PDT, dozens of customers have reported slow performance and frequent timeouts on API calls. Dashboards indicate that about 27% of API calls are exceeding SLO targets.
Conditions: Responders have determined that the slow and failing API calls are all related to the users database table. It appears that a database schema change that rolled out at about 10:00 PDT missed a critical index on the users table, which is making database calls that access or update that table much slower than expected.
Actions: @lynn from the Databases team is preparing a further database schema change PR to redefine and regenerate the missing index; @ravi is standing by to review the PR as soon as it is ready. @sami from Customer Care has published a banner on the “report a problem” web page to let customers know that we’re aware of the problem, and @jamie from Developer Relations is sending an email to Tier 1 customers who use the API, informing them of the problem and that we’re working on a fix.
Needs: @ravi from the Databases team needs to finish reviewing and approving the database schema change PR with the fix. Then, we need to deploy it and wait while the index is rebuilt. We currently estimate that the rebuild will take approximately 2 hours, but we won’t know for sure until it is underway and we can see how fast it is proceeding.
Expect the next update in 1 hour, or sooner if circumstances warrant. If you have any questions or concerns, please bring them up in the incident channel (#i-5150-api-timeouts), or DM them to the Incident Commander (currently @brent).
By following these best practices, you can create SitReps that are clear, concise, and contribute to effective incident management.
—–
IT incidents can be costly, impacting both customers (through lost revenue and damaged reputation) and staff (with reduced productivity, decreased morale, and increased turnover). Proactive preparation and continuous learning from incidents are crucial. As an expert incident management consultant, advisor, and coach, I can help your organization develop these critical skills and avoid costly mistakes. Contact me today to learn more.
Security and availability incidents share many similarities, but also have key differences that, if understood and addressed, can significantly improve your incident management.
Similarities:
Both types of incidents require forming and coordinating ad hoc teams of responders.
Challenges include identifying the right responders, fostering effective collaboration among them, and keeping stakeholders informed.
Key Differences:
Duration: Availability incidents typically get resolved in hours (e.g., when I led incident management at Slack, about half of our availability incidents were resolved in under 2 hours, and two-thirds in under 4 hours). Security incidents often last days or weeks.
Secrecy: Availability incidents are usually handled openly within a company. Security incidents are often handled secretly, especially if an intruder might be monitoring communications.
Ideally, a single incident management process and unified tools should be used for both availability and security incidents. You don’t want your engineering teams to have to learn two different processes and tool sets for the different types of incidents. This requires adapting the process and tools to accommodate the differences, though:
Extended Duration of Security Incidents
Develop robust handoff procedures for responsibilities and information as responders transition in and out over time.
Refine methods for providing ongoing progress updates, as a single resolution summary at the end of the incident (days or weeks later) will not suffice.
Secrecy/Visibility of Security Incidents:
Enable incident commanders and responders to collaborate securely while keeping details (or even the incident’s existence) confidential from those without a legitimate need to know.
This may involve using restricted communication channels (e.g., a private Slack channel instead of a public one) and tightly controlling access to documents (e.g., access permissions in Google Drive/Docs/Sheets).
One tip: consider sharing “placeholder” information about security incidents publicly within the company. Even if details are secret, acknowledge the incident’s existence and direct inquiries to a point of contact (e.g., the Incident Commander or a senior security manager). This can be posted in the incident’s automatically created public Slack channel while responders work in a parallel private channel. However, this approach is not suitable if the mere existence of the incident must be kept secret (e.g., when chasing an active intruder who may have access to your Slack workspace). For most security incidents, even simply acknowledging the incident can address many concerns and improve security engagement within the company.
Finally, even if the security incident needs to be kept secret while it is underway, don’t assume that it must be kept secret forever. There might be valuable lessons to learn from broader visibility once the incident is over.
—
IT incidents can be incredibly costly for you with both your customers (resulting in lost revenue, missed sales, and damaged reputation) and your staff (resulting in decreased productivity, reduced morale, and increased turnover). That’s why it’s crucial to be proactive, prepare for future incidents, and learn as much as you can from every incident. As an expert incident management consultant, advisor, and coach, I can guide your organization in developing these critical skills and help you avoid expensive mistakes. Contact me today to learn how I can help your team.
Delegation is a critical skill for incident commanders. Have you ever considered how tasks are delegated during an incident, though, and what impact that can have on the incident? Do the incident commander and the responders share a common understanding of exactly what is being delegated, how much freedom to act the responders have, and how much responsibility the responders are expected to take?
Someone recently shared a neat one-pager on Levels of Delegation with me, which summarizes 10 different levels of delegation, ranging from 1 to 10 in terms of the amount of freedom and responsibility being delegated. For incident management, key levels include:
Level 1: “Wait to be told,” “Do exactly what I say,” or “Follow instructions precisely.”
Level 3: “Look into this and tell me the situation. We’ll decide together.”
Level 5: “Give me your analysis and recommendation. I’ll decide if/when to proceed.”
Level 6: “Decide and let me know your decision, and wait for my go-ahead.”
Level 7: “Decide and let me know your decision, then go ahead unless I say not to.”
Level 8: “Decide and take action; let me know what you did.”
Level 10: “Decide where action needs to be taken and manage the situation accordingly. It’s your responsibility now.”
Incident commanders typically aim for Level 5 or 6, expecting responders to investigate, determine actions, and then report back before implementation to ensure coordination. The problem is that responders might be operating under different assumptions (e.g., Level 3 for joint decision-making, or Level 8 for independent action).
To avoid mismatched expectations, be explicit when delegating. Clearly state the expected level of autonomy, for example, “look, but don’t change anything without coordinating with me first.”
Also, be explicit about the expected timeframe. Establish clear timelines, even if it’s just for an update like “we’re still working on it; we expect to have more info in 20 minutes.”
Finally, be clear about the required level of accuracy or certainty. A rough estimate might take just minutes, whereas determining a precise figure could take hours. If “between 5,000 and 10,000” is a good enough answer for now, say so.
Responders also share responsibility for clear communication. If you’re uncertain about the delegation level, timeframe, or desired accuracy, ask for clarification.
Effective incident management depends on clear delegation and open communication between incident commanders and responders.
Incidents are costly, impacting both customer relationships (leading to lost revenue, missed sales, and reputational damage) and staff (through decreased productivity, reduced morale, and increased turnover). It is essential to be proactive by both preparing for future incidents and thoroughly learn from each one that occurs. As an expert incident management consultant, advisor, and coach, I help companies develop these vital capabilities. Contact me today to discover how I can help your team.
Recent Comments