Context
Cloudflare sits between users and millions of websites, providing security, CDN, and DDoS protection. A huge portion of the internet flows through Cloudflare’s edge network.
What Happened
On July 2, 2019, a new firewall rule was deployed globally across Cloudflare’s network.
It contained a seemingly harmless regular expression designed to detect malicious traffic.
But hidden inside it was a performance disaster.
The regex triggered catastrophic backtracking, causing CPUs on Cloudflare’s edge servers to spike to nearly 100% utilization almost instantly.
Within seconds, servers across the globe began choking.
Latency skyrocketed. Requests failed. Entire regions of the internet slowed to a crawl.
Sites that depended on Cloudflare including Discord, Shopify, and many others became inaccessible.
All because of a single line of code.
Source: https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019/
Root Cause
A poorly optimized regular expression caused exponential CPU consumption (catastrophic backtracking), overwhelming edge servers globally.
Impact
Massive global slowdown and outages
Thousands of websites affected simultaneously
Traffic drops across major platforms
Widespread internet disruption for ~30 minutes
Fix
Cloudflare engineers quickly identified the offending rule and rolled it back.
Traffic normalized almost immediately once CPU load dropped.
Lessons Learned
- Even small code changes can have global impact at scale
- Performance bugs can be as dangerous as functional bugs
- Global deployments amplify risk instantly
- Edge systems must be protected from unbounded resource usage
Prevention
- Introduce staged rollouts instead of global pushes
- Add performance testing for regex and heavy computations
- Implement automatic safeguards for CPU spikes
- Use canary deployments before full rollout