
CrowdStrike and the Global Domino Effect of July 19, 2024
The CrowdStrike Case: A Discussion on the Causes.
On July 19, 2024, CrowdStrike released an update for their Falcon sensor, which caused a massive crash on numerous Windows computers, leading to global disruptions in critical sectors such as aviation, media, banking, and healthcare. The incident was caused by a logic error in an updated configuration file, which triggered an infinite reboot loop on affected devices (CrowdStrike) (Morningstar) (Wikipedia).
About This Article
This article originated from a discussion and social interaction on the professional social network LinkedIn (https://linkedin.com). What is reported here summarizes and extends the content, opinions, and viewpoints discussed among the various people mentioned at the time of the events, in the days following the global outage of the commercial product CrowdStrike Falcon (July 19, 2024).
Responsibility and Professional Ethics
Roberto Beneduci, CEO of CoreTech, expressed concern about the blame erroneously attributed to Microsoft instead of the actual responsible party, CrowdStrike. Beneduci compared the situation to an excavator cutting internet cables while the media blames the internet service provider for the outage. He emphasized the importance of correctly attributing responsibility and ensuring that software products do not cause harm.
Technical and Product Perspectives
Ivan Dorna, CEO of Anthilla, expanded the discussion by explaining that the problem does not primarily reside in the operating system, while acknowledging its current state and limitations, but rather in the awareness and choices made when setting up an infrastructure that must deliver a service, and in the provision of a third-party service that serves that infrastructure. This stems from the competence, ethics, and professionalism of those making the choices, but also from the constraints (often imposed and not always understood) related to the activities to be performed.
He highlighted the importance of technical competence to quickly resolve problems, noting the delay — not so much in spreading the workaround, which is essentially the standard one for driver issues on Windows, but in the fact that it seems strange that this procedure was not applied earlier, as it should be known and foundational to technical knowledge for Microsoft systems.
He also noted that similar incidents could occur on Linux under analogous circumstances.
There is never a single right or correct answer by definition, which is why the actual risk of all software types must be analyzed. This software type has been “selected” and deployed as an enabling element (i.e., an essential component to validate a platform from a security standpoint and obtain insurance coverage), further tying the burden of service delivery to one or more commercial products rather than to an actual independent risk assessment that would evaluate technologies and product maintenance policies.
Ivan Dorna discussed this point, namely the identification of CrowdStrike as an “Enabling Element” both as a software type and as a well-known brand, to activate insurance coverage globally, without adequately considering the real risks.
This led to a situation where companies imposed a product for its perceived benefits without evaluating the impact in case of malfunction, which runs counter to actual risk factor assessment. Not allowing a deeper investigation of the real situation.
“The argument about mandatory requirements for insurance coverage is absurd. Choosing the system for the enabling factor introduced a series of additional risks.”
This is true both from the perspective of the “enabling factor” for insurance and for certain constraints and obligations deriving from best practices, as will also be the case for NIS2 and DORA.
The notoriety of a product, as has been the case in the past for other brands and products, can no longer be sufficient to recognize them as “enabling elements.” As with VMware, not only regarding the directional changes of 2023 and 2024, but also in light of the policy and strategy changes made from 2010 to today.
Economic and Security Considerations
Andrea Campiotti mentioned that there have also been issues with Azure AD and suggested there could be shared blame between CrowdStrike and Microsoft.
Giulio Covassi emphasized the importance of quality over new features, highlighting how untested software can cause significant damage. He suggested that a platform engineering approach could mitigate these risks.
Francesco Vollero reiterated that if the operating system did not allow the development of kernel modules running with elevated privileges, the problem would have been less severe. He criticized the current organizational structure that often leads to errors due to lack of adequate testing.
Luca Simoncini called for a detailed Root Cause Analysis from CrowdStrike to learn from the mistakes made.
Matei Busui suggested a procedure to classify critical and non-critical services and stressed the importance of integration testing and rollback procedures.
Matei Busui: “If you run critical services, you must ensure they don’t stop for any reason during an update. Additionally, it’s fundamental to have integration tests that validate everything works as expected after an update.”
Ivan Dorna: “Applying this to the real world means working to make it possible to roll back to the previous version in every case, within minutes.”
Discussion on the 1% and Quantity Relationships
Nicola Vanin reported that approximately 8.5 million devices, less than 1% of Windows computers globally, were affected by the incident.
He observed that, despite the low percentage, the chaos generated was significant.
Massimo Biagioli commented that speaking in terms of percentages can be misleading, and that it is more relevant to consider absolute numbers to understand the real impact.
Ivan Dorna:
“The percentage must be put in relation to the context and the specific problem. 1% of blocked computers, equivalent to 8.5 million devices, has a significant global impact, and this data must not be overlooked.”
Relating the 1% value, not so much to the absolute value but as a threshold value, for what it represented as a problem on July 19, 2024, is now important and must be taken into consideration.
Nowadays, both due to the spread of software and its use in complex installations that deliver services, perhaps also in the cloud, the impact of a problem on even a secondary software component that must be delivered and kept updated for a whole series of reasons cannot be neglected.
Considerations on Procedures and Service Management
As of today, July 21, 2024, the management of program execution as services in operating systems is not optimal — in fact, it is very primitive. Since these must be updated frequently for compliance with legal requirements and corporate best practices, especially for services with administrative and system privileges, it is fundamental to improve the robustness of this software and its management system.
The problem experienced with CrowdStrike is due to three main reasons:
- The update on a Stable/Production channel was not sufficiently tested, especially for a component that operates at a privileged level.
- A product + service was used as an “enabling element” for insurance coverage, without an actual risk assessment and recovery procedures.
- Today, with the spread of ITIL and certification frameworks for the compartmentalization of roles and responsibilities, it seems that technicians’ skills and ability to act are limited. This raises the question: does it really make more sense today to assign blame rather than keep things running and restore them?
Conclusion
This incident highlighted the importance of conscious and competent management of software updates and corporate responsibilities. The discussion among IT experts underscores the need for a holistic approach to cybersecurity, which includes not only software quality, but also the responsibility of those who develop and manage it.
Sources of the LinkedIn discussions:


