Context
By 2017, Amazon S3 had quietly become the backbone of the modern internet. From startups to global giants, countless applications depended on it to serve images, videos, APIs, and critical data. Most users didn’t know it - but the internet was standing on AWS.
What Happened
On the afternoon of February 28, 2017, inside an AWS operations environment, a site reliability engineer initiated what should have been a routine debugging procedure.
The goal was simple: temporarily remove a small number of servers from service.
But something went wrong.
A command was entered incorrectly- removing far more servers than intended.
Within seconds, critical subsystems inside S3—particularly the index and placement systems responsible for locating and serving data—began to collapse.
Requests started failing. Then retry storms began. Then dashboards went red.
Across the world, engineers watched in disbelief as their services began to break - Slack wouldn’t load, Trello froze, Quora vanished.
For hours, large parts of the internet simply… stopped working.
🔗 Source: https://aws.amazon.com/message/41926/
Root Cause
A mistyped command executed without sufficient safeguards, combined with deeply coupled internal systems and recovery paths that had not been exercised at full scale.
Impact
Major global services went offline simultaneously
Thousands of businesses disrupted
Developers worldwide locked out of their own systems
Estimated economic damage exceeding $150 million
Fix
AWS engineers initiated a controlled restart of the affected subsystems.
But recovery wasn’t instant—because the systems had not been restarted at this scale in years, bringing them back online took significantly longer than expected.
Gradually, over several hours, S3 stabilized and the internet began to recover.
Lessons Learned
- Even the smallest human mistake can cascade into global failure
- Critical infrastructure must assume operator error
- Rarely used recovery procedures become liabilities
- Hidden dependencies amplify failure impact
Prevention
- Introduce strict safeguards for destructive commands
- Limit blast radius through system isolation
- Continuously rehearse full-scale recovery scenarios
- Design systems to degrade gracefully instead of collapsing