OnlyTech.boo

The us-east-1 Domino Collapse

April 6, 2026$100000k estimated cost

Context

Amazon Web Services us-east-1 is the most critical AWS region, powering a massive portion of global internet infrastructure from startups to Fortune 500 companies.

What Happened

On December 7, 2021, at around 10:30 AM ET, AWS engineers were working on scaling capacity for a service within the us-east-1 region. A change was made to the internal network controlling communication between services. Almost immediately, something went wrong. A surge of unexpected traffic flooded the internal network, overwhelming key systems responsible for service coordination. APIs began failing. Monitoring systems went blind. Then came the cascade. Services that depended on these internal systems DynamoDB, EC2, Lambda started degrading or failing entirely. Soon, major platforms like Netflix, Disney+, and even Amazon’s own services began experiencing outages. For hours, one of the most critical regions of the cloud was partially crippled. Source: https://aws.amazon.com/message/12721/

Root Cause

A network configuration change triggered a traffic surge that overwhelmed internal service communication systems, causing cascading failures across dependent services.

Impact

Major outages across thousands of companies Streaming platforms, apps, and IoT devices disrupted Amazon warehouse operations impacted Multi-hour degradation of core AWS services

Fix

AWS engineers throttled traffic and rebalanced network load. Gradually, dependent services recovered as internal communication stabilized.

Lessons Learned

  • Internal service dependencies can create hidden single points of failure
  • Traffic surges inside systems are as dangerous as external attacks
  • Observability systems must remain functional during failures
  • Regional concentration increases systemic risk

Prevention

  • Improve isolation between internal service components
  • Add rate limiting and safeguards for internal traffic
  • Strengthen multi-region architectures
  • Ensure monitoring systems are independent and resilient

Similar incidents

PocketOS operated as a SaaS platform for car rental businesses, running on cloud infrastructure with shared storage volumes across staging and production. An AI coding agent inside Cursor, powered by a model from Anthropic, was granted execution capabilities within this environment. The system served real customers with live transactional data. A small engineering team managed infrastructure, application logic, and deployments. Stakeholders included rental operators, end users, developers, and infrastructure providers such as Railway.

Comments

Oldest first.

Loading comments…