The full post has depth I couldn't fit here: consistent hashing internals, Multi-Paxos vs. leaderless tradeoffs, interactive timelines, and what all of it means for your schema design.
Link in the first reply.
2012: DynamoDB. The 2022 USENIX ATC paper from Amazon engineers said it directly:
"These architectural discussions culminated in Amazon DynamoDB...shared most of the name of the previous Dynamo system but little of its architecture."
The partition is still the unit of scale, inherited directly from Dynamo. You can't join across partitions, so you denormalize everything upfront.
The API is small: get by key, query by sort key range, scan. That tight surface is what makes the performance guarantees possible.
AWS also built SimpleDB around this time. Managed, clean API. You didn't run it yourself.
Hard cap: 10 GB per domain. Every attribute auto-indexed, making writes expensive. Worked for startups. Broke when you needed real data volumes.
By the late 2000s, Dynamo was everywhere inside Amazon. Teams tolerated it more than loved it.
You couldn't just call an API. You ran a cluster. You had to understand consistent hashing well enough to debug production incidents at 2am.
Keys mapped to a ring via consistent hashing. New node joins, only its slice of keys move. Virtual nodes smoothed uneven distributions. Gossip protocol handled topology β no central metadata service. Every node learned about failures through heartbeats.
π§΅ In 2004, Amazon's shopping carts went down during the holiday rush. Databases saturated. Real revenue, gone per second.
That failure started an 8-year path to DynamoDB. The path was weirder than most people realize:
Two replicas, both accepting writes during a network partition. Now you have conflicts. Dynamo's solution: vector clocks. Every write tagged with logical timestamps per node.
When versions diverged, the application resolved it. The shopping cart team merged carts by union.
Dynamo's premise was blunt: if your shopping cart rejects a write because a replica is unavailable, you lose a sale.
So they made it always-writeable. Every replica could accept writes. No leader required. Sloppy quorum.
Amazon's response was Dynamo. Built internally. Published at SOSP 2007. Purpose-built for one thing: state management in a service-oriented architecture at Amazon's scale.
Dynamo was never a public service. You couldn't buy it. Internal only.
At that scale, schema changes meant maintenance windows. Want to add a column to the shopping cart table? Notify every team. Plan downtime. Across dozens of fast-growing services.
Oracle was a shared dependency. Shared dependencies become bottlenecks.
Amazon ran on Oracle. The problem: scaling limits. Every connection ate server memory. Hit the limit and you chose between rejecting clients or buying more Oracle licenses.
Both options were terrible at scale.
DynamoDB Beginner trap: using `Scan` to find items, then filtering in your application code.
DynamoDB charges you for every item it reads, even the ones your filter throws away.
If you're filtering out 90% of results, you're paying 10x what you should.