The short answer
Quick answer: The largest technology companies built their own databases because, at the time, nothing they could buy survived their workload. General-purpose databases of the early 2000s were designed to run on one powerful machine. Companies such as Google and Amazon needed to spread data across thousands of ordinary machines, stay available while some of those machines were failing, and do it at a cost that licences per server made impossible. So they wrote systems tuned to their own access patterns and trade-offs. Their published designs became the basis of most modern distributed databases, which is the main reason almost nobody else needs to build one today.
The problem they faced
A traditional relational database grows by scaling up: a bigger server, more memory, faster disks. That works until the largest available machine is not big enough, and long before then it becomes very expensive.
The alternative is scaling out across many machines. In the early 2000s the mainstream databases did not do this by themselves. Teams split the data by hand, a technique called sharding, and built their own tooling to route queries, move data and recover from failures. At a certain size that tooling is a database, written badly and by accident.
Other pressures pushed in the same direction:
- Failure is constant. With thousands of machines, something is always broken. The system must keep working through it, without a person intervening.
- Global users need data in several regions, close to them.
- Cost. Commercial licences priced per processor become enormous across a large fleet.
- Narrow workloads. A shopping basket or a web index needs a few simple operations at huge volume, not the full generality of SQL.
The systems that changed the field
Google Bigtable (2006)
Google needed to store the web index and data for products such as Maps and Earth. The Bigtable paper describes a "sparse, distributed, persistent multidimensional sorted map" designed for petabytes across thousands of commodity servers. Its storage design, based on sorted, immutable files that are merged in the background, made the LSM tree popular. See B-trees vs LSM trees.
Open-source descendants: HBase, and part of Cassandra's design.
Amazon Dynamo (2007)
Amazon's shopping basket had to accept writes always, even during failures, because a rejected "add to basket" is lost revenue. The Dynamo paper describes a key-value store that chooses availability over strict consistency. It combined consistent hashing, replication with tunable quorums, version tracking and a gossip protocol, and accepted eventual consistency as the price.
The paper's Dynamo is an internal system. The cloud service DynamoDB is a later, different product that took its name and some of its ideas.
Open-source descendants: Cassandra, Riak, Voldemort.
Cassandra (2008)
Built at Facebook for inbox search and then open-sourced, Cassandra combined Dynamo's distribution model with Bigtable's data model. It is now an Apache project used widely outside the company that created it.
Google Spanner (2012)
The first wave gave up transactions and SQL to gain scale, and application developers paid for it in complexity. Spanner was Google's answer: a database that is distributed across the globe and offers strongly consistent transactions and SQL. The Spanner paper describes how it uses TrueTime, an API backed by GPS receivers and atomic clocks that reports time with a known bound of uncertainty, to order transactions across data centres. See clocks in distributed systems.
Systems inspired by it: CockroachDB, YugabyteDB, TiDB.
Summary
| System | Built by | Built for | Key choice |
|---|---|---|---|
| Bigtable | Huge, sparse tables | Scale and throughput; LSM storage | |
| Dynamo | Amazon | Always-writable key-value data | Availability over consistency |
| Cassandra | Write-heavy data across many nodes | Dynamo's distribution, Bigtable's model | |
| Spanner | Global data needing transactions | Strong consistency using synchronised clocks |
Many other companies did the same in their own niches: stores for social graphs, time-series metrics, analytics and logs.
Why custom designs paid off
Different trade-offs
A general-purpose database must be acceptable at everything. A purpose-built one can be excellent at one thing. The CAP theorem says a distributed system must choose how to behave when the network splits; Dynamo and Spanner made opposite choices, each correct for its workload.
Fit to the access pattern
If every query is "fetch this key" or "scan this range", you can drop joins, the query planner and much else, and gain speed and simplicity.
Economics at scale
For a company running hundreds of thousands of servers, a few percent of efficiency is worth more than the salaries of the team that finds it. The same sum is a rounding error for a small company.
Control
Owning the system means being able to diagnose and fix any problem, integrate with internal infrastructure, and avoid dependence on a vendor's priorities.
People
These companies could hire and retain large teams of distributed systems specialists, and keep them on the problem for a decade.
The real cost
Building a database is among the hardest tasks in software.
- Correctness is brutally difficult. Losing or corrupting data is unforgivable, and the bugs appear only under rare combinations of failure. See why distributed systems are hard.
- It takes years before a new database is trustworthy.
- Everything around it must be built too: backup, monitoring, migration tools, client libraries, documentation.
- Maintenance never ends, and new hires must learn a system that exists nowhere else.
Even the giants do not do it casually. They run plenty of MySQL and PostgreSQL as well, often with heavy sharding layers on top.
Why you almost certainly should not
The conditions that justified those systems no longer apply to most organisations.
- The ideas are available off the shelf. Open-source and cloud databases now offer horizontal scaling, replication and automatic failover.
- Hardware is far more capable. A single modern server with fast SSDs and hundreds of gigabytes of RAM handles workloads that needed a cluster in 2005.
- Most scale problems are not database-engine problems. They are missing indexes, inefficient queries, and absent caching. See how database indexes work.
A sensible order of escalation:
- One well-tuned relational database.
- Add caching and read replicas. See how database replication works.
- Use a specialised engine for a specialised need: search, analytics, time series.
- Adopt a distributed database.
- Shard.
- Only if you have a workload that truly nothing serves, and the team to sustain it, build.
Very few organisations reach step 6. For the choice between existing families, see SQL vs NoSQL.
What everyone gained
The most valuable output of this work was not the software but the papers. By publishing their designs, these companies taught the industry how to build distributed storage. An entire generation of open-source databases followed, and the cloud providers turned the same ideas into managed services anyone can rent by the hour.
Frequently asked questions
Why did Google build Bigtable?
To store and serve very large structured datasets, such as its web index, across thousands of commodity machines, which existing databases could not do.
Is DynamoDB the same as Dynamo?
No. Dynamo is the internal system described in Amazon's 2007 paper. DynamoDB is a later managed cloud service that shares the name and some concepts.
Should a startup build its own database?
Almost never. Existing open-source and managed databases cover nearly every need, and building one is a multi-year effort with a high risk of data loss.
Do big companies still use MySQL and PostgreSQL?
Yes, widely, often alongside their custom systems and with additional layers for sharding.
Conclusion
Big tech companies built databases because their scale arrived before the tools did. Each system was a specific answer to a specific workload, with deliberate trade-offs. Their lasting contribution is that the answers were published and reimplemented, so the rest of us can choose a database instead of writing one.
Related articles
- SQL vs NoSQL: When to Use Which (Beyond the Hype)
- Sharding Explained: How Databases Scale Beyond One Machine
- The CAP Theorem Explained With Real Examples
- B-Trees vs LSM Trees: Why Databases Store Data Differently
