OnlyTech.boo

The Day a Single Command Broke the Internet — AWS S3 Outage (2017)

April 6, 2026$15000k estimated cost

Context

By 2017, Amazon S3 had quietly become the backbone of the modern internet. From startups to global giants, countless applications depended on it to serve images, videos, APIs, and critical data. Most users didn’t know it - but the internet was standing on AWS.

What Happened

On the afternoon of February 28, 2017, inside an AWS operations environment, a site reliability engineer initiated what should have been a routine debugging procedure. The goal was simple: temporarily remove a small number of servers from service. But something went wrong. A command was entered incorrectly- removing far more servers than intended. Within seconds, critical subsystems inside S3—particularly the index and placement systems responsible for locating and serving data—began to collapse. Requests started failing. Then retry storms began. Then dashboards went red. Across the world, engineers watched in disbelief as their services began to break - Slack wouldn’t load, Trello froze, Quora vanished. For hours, large parts of the internet simply… stopped working. 🔗 Source: https://aws.amazon.com/message/41926/

Root Cause

A mistyped command executed without sufficient safeguards, combined with deeply coupled internal systems and recovery paths that had not been exercised at full scale.

Impact

Major global services went offline simultaneously Thousands of businesses disrupted Developers worldwide locked out of their own systems Estimated economic damage exceeding $150 million

Fix

AWS engineers initiated a controlled restart of the affected subsystems. But recovery wasn’t instant—because the systems had not been restarted at this scale in years, bringing them back online took significantly longer than expected. Gradually, over several hours, S3 stabilized and the internet began to recover.

Lessons Learned

  • Even the smallest human mistake can cascade into global failure
  • Critical infrastructure must assume operator error
  • Rarely used recovery procedures become liabilities
  • Hidden dependencies amplify failure impact

Prevention

  • Introduce strict safeguards for destructive commands
  • Limit blast radius through system isolation
  • Continuously rehearse full-scale recovery scenarios
  • Design systems to degrade gracefully instead of collapsing

Similar incidents

PocketOS operated as a SaaS platform for car rental businesses, running on cloud infrastructure with shared storage volumes across staging and production. An AI coding agent inside Cursor, powered by a model from Anthropic, was granted execution capabilities within this environment. The system served real customers with live transactional data. A small engineering team managed infrastructure, application logic, and deployments. Stakeholders included rental operators, end users, developers, and infrastructure providers such as Railway.

Comments

Oldest first.

Loading comments…
The Day a Single Command Broke the Internet — AWS S3 Outage (2017) - OnlyTech.boo