In the early hours of July 19th, companies in Australia running Microsoft’s Windows operating system began reporting Blue Screens of Death (BSODs) on their devices. Following this, similar disruptions were reported from the UK, India, Germany, the Netherlands, and the United States. Notably, TV station Sky News went offline, and United, Delta, and American Airlines issued a “global ground stop” on all flights. This led to major U.S. airlines grounding flights and causing global delays. Also affected were industries like healthcare, banking, agriculture, and even government.
The cause of these widespread Windows outages was traced to a software update from the cybersecurity company CrowdStrike. No malicious cyberattack had occurred– the incidents stemmed from a misconfigured or corrupted update pushed out to CrowdStrike’s customers.
When events like this occur, it can be easy to point fingers. However, this type of adverse event could’ve happened to just about any software, service, or other piece of technology you use today– and without any proactive planning, the effects can be devastating.
Be prepared
Technology outages are common, due in part to change management issues, maintenance neglect, cyberattacks by bad actors, or simply nature and the unpredictability of the world today. This shouldn’t dissuade you from making changes– it should in fact encourage you to do just that.
1. Invest in Disaster Recovery
Disaster Recovery (DR) is essential insurance for navigating unstable times. The key is to implement and test while things are stable.
A great DR practice utilizes several key factors to mitigate risk: maintenance, backups, diversions of traffic, and tests to critical systems. It works to shorten downtime, getting systems back up and running as quickly as possible and ensuring that crucial data isn’t leaked or lost in the process. Disaster Recovery encapsulates actions like:
- Replicating servers in disparate locations
- Rerouting traffic to unaffected infrastructure
- Designating and training crisis roles for team members
- Testing systems to ensure feasibility in an emergency
- Keeping generators on deck to mitigate power loss
- Disaster Recovery as a Service (DRaaS) implementation
DRaaS is a service model that reduces the cost of owning and maintaining a secondary data center. DRaaS providers host your critical information and applications securely in the cloud, so that if physical hardware and servers are destroyed, you may still access the affected assets and run key applications.
Whether you choose to go with the As-a-Service model or go traditional with a second data center, the important thing is that you invest in either solution before you need it. Once data is gone, it’s gone– and downtime can be extremely expensive.
Having a disaster recovery plan ensures that your organization can get back online quickly, keep emergency costs as low as possible, and protect customer data– three factors that hold the most weight in defining your reputation and empowering continued operations in a crisis.
2. Communicate in a crisis
As we highlighted earlier, CrowdStrike’s response to this emergency was thorough, swift, and was commended by many. The following comes from CIODive:
Mea culpas are sparse in the cybersecurity industry. CrowdStrike accepted an usual level of accountability of its own volition — and without condition.
“Kurtz’s quick apology for a defective software update is rare in cybersecurity — I can’t think of any other case — but reflects a growing trend of corporate accountability,” Mauricio Sanchez, senior director of enterprise security and networking research at Dell’Oro Group, said in an email.
CrowdStrike customers, affected members of the public and industry observers alike appreciated the honesty and rapidness with which they responded to the incident. But, it wasn’t just the apology that made a difference. It was the continued communication–in the form of information sharing, regular updates, and resource creation– that cemented CrowdStrike’s trustworthy reputation and led to quicker resolutions.
Their remediation included sharing a technical explanation of the issue, pushing out a self-guided instructional video that allowed users to solve the problem on their own, and staying ahead of the curve on potential malicious activity from bad actors leveraging the outage. To date, they continue to release updates, reports, and more on their dedicated Falcon Content Update Remediation and Guidance Hub.
Communicating internally is another key to a successful remediation. Within hours of initial reports, CrowdStrike was able to pinpoint the source of the issue, down to the exact error in the exact program of concern. This demonstrates an ability many businesses should aspire to: the art of communication amongst the ranks.
Prepare a chain of command, a contact list, and strategies for tracing the source of an incident among your internal team. Ensure that everyone is kept in the loop on potential risks, including:
- Recent phishing attempts or other issues experienced by team members
- Upcoming software updates
- New product or service implementations
- Changes to IT infrastructure
- Third-party breaches in programs used by your organization
- Events that raise risk (holidays, natural events, and other external happenings)
They say a team is only as strong as its weakest member. When everyone is aware of risk, everyone can contribute to a high level of cyber hygiene and plan ahead for potential interruptions. Communicate clearly and openly the risks, expectations, plans, and response procedures put in place, and allow your team to grow stronger with that knowledge.
3. Practice for the unexpected
When one has the most thorough plans, it seems inevitable that something goes awry. Or at least, it might feel that way sometimes– you plan an entire vacation and your outbound flight gets delayed, you order dinner ahead and the kitchen’s backed up, you plan a beach day and wind up in a surprise sunshower. It happens to everyone.
Unfortunately, it’s only after an inconvenience that we are able to do better next time: packing a pillow for the airport, keeping a quick snack on hand, or placing an umbrella in the beach bag. Though in your daily life, these inconveniences may be mostly inconsequential, your organization doesn’t have that privilege. Unfortunately, it only takes one incident to bankrupt a company, or at least to have devastating effects on their reputation.
Since learning from organic experience isn’t always an option, practice is crucial. HHS defines an Incident Response Plan as “a written document that helps your organization before, during, and after a security incident.” An all-encompassing IRP can include things like:
- Team roles and responsibilities
- Procedures for analysis, detection, and mitigation of an active threat
- A summary of service providers and internal resources to utilize
- Communication instructions
The best way to learn is by doing. When you’ve drawn up your IRP, you should practice it, too.
DR testing is a critical factor in ensuring the feasibility of your plan and the preparedness of your team to follow it. A disaster recovery test can help “determine if a DR plan can work and meet an organization’s predetermined RPO/RTO requirements. It also provides feedback to enterprises so they can amend their DR plan should any unexpected issues arise. (TechTarget)”
You can conduct DR testing through several methods, including tabletop exercises and simulation, but ideally the commonality is hands-on participation from team members who will carry responsibilities in a real emergency. Strengthening the muscles and discovering potential pitfalls before the big day will prove invaluable should you ever find yourself in a crisis.
The bottom line
No one is exempt from errors, crises, and emergencies– not a single provider or end-user is guaranteed not to be struck someday. The differentiator lies in being prepared: when disaster strikes, will you know how to respond?
UPSTACK firmly believes in the power of proactivity. Our experts have extensive understanding of the risks associated with technology, and we maintain the most up-to-date preparation methodology possible. Suppliers are equally vetted for their mitigation and response capabilities, and we choose to work with disaster recovery specialists with a proven track record of success.
Ready to protect your business from whatever comes next? Speak with our experts to learn how we can be a lifeline if disaster strikes.



