Context
Amazon Web Services us-east-1 is the most critical AWS region, powering a massive portion of global internet infrastructure from startups to Fortune 500 companies.
What Happened
On December 7, 2021, at around 10:30 AM ET, AWS engineers were working on scaling capacity for a service within the us-east-1 region.
A change was made to the internal network controlling communication between services.
Almost immediately, something went wrong.
A surge of unexpected traffic flooded the internal network, overwhelming key systems responsible for service coordination.
APIs began failing. Monitoring systems went blind.
Then came the cascade.
Services that depended on these internal systems DynamoDB, EC2, Lambda started degrading or failing entirely.
Soon, major platforms like Netflix, Disney+, and even Amazon’s own services began experiencing outages.
For hours, one of the most critical regions of the cloud was partially crippled.
Source: https://aws.amazon.com/message/12721/
Root Cause
A network configuration change triggered a traffic surge that overwhelmed internal service communication systems, causing cascading failures across dependent services.
Impact
Major outages across thousands of companies
Streaming platforms, apps, and IoT devices disrupted
Amazon warehouse operations impacted
Multi-hour degradation of core AWS services
Fix
AWS engineers throttled traffic and rebalanced network load.
Gradually, dependent services recovered as internal communication stabilized.
Lessons Learned
- Internal service dependencies can create hidden single points of failure
- Traffic surges inside systems are as dangerous as external attacks
- Observability systems must remain functional during failures
- Regional concentration increases systemic risk
Prevention
- Improve isolation between internal service components
- Add rate limiting and safeguards for internal traffic
- Strengthen multi-region architectures
- Ensure monitoring systems are independent and resilient