Search

Why Clocks Can't Be Trusted in Distributed Systems

The short answer

Quick answer: Every computer has its own clock, and no two agree exactly. Clocks drift at different rates, are corrected in jumps by synchronisation, and can even move backwards. So if machine A stamps an event 10:00:00.005 and machine B stamps another 10:00:00.003, you cannot conclude that B's event happened first. Distributed systems therefore avoid using wall-clock time to order events. They use logical clocks, such as Lamport timestamps and vector clocks, which capture cause and effect, or special infrastructure such as Google's TrueTime, which reports time together with its uncertainty.

Why computer clocks are wrong

A computer keeps time with a quartz crystal oscillator. Crystals are cheap and slightly inaccurate, and their frequency changes with temperature and age. A typical clock drifts by tens of parts per million. That sounds tiny, but it adds up to seconds per day.

To correct drift, machines synchronise with time servers using NTP (Network Time Protocol). NTP estimates the network delay to the server and adjusts the local clock. But:

  • The estimate assumes the delay is the same in both directions, which is often untrue.
  • Over the public internet, accuracy is typically in the range of milliseconds to tens of milliseconds.
  • A misconfigured or unreachable NTP server leaves the clock drifting.
  • Virtual machines can be paused, and their clocks may jump when they resume.

Two terms are worth separating:

  • Drift: how fast a clock gains or loses time.
  • Skew: the difference between two clocks at a given moment.

Clocks can jump

When NTP finds the clock is wrong, it either slews it (runs it slightly faster or slower for a while) or steps it (sets it directly). A step can move time backwards. Add leap seconds, which occasionally insert an extra second into the day, and you have a clock that might repeat a second or stall.

Some large providers "smear" leap seconds across many hours to avoid a sudden jump.

Two kinds of clock

Operating systems offer two different clocks, and mixing them up is a common bug.

Wall clock (time of day)Monotonic clock
Tells youThe current date and timeTime elapsed since an arbitrary point
Can jump or go backwardsYesNo
Comparable across machinesRoughlyNot at all
Use forShowing times to people, timestamps in logsMeasuring durations, timeouts, rate limits

Measuring how long something took by subtracting two wall-clock readings can give a negative result if the clock was adjusted in between. Always use a monotonic clock for durations.

import time
start = time.monotonic()
do_work()
elapsed = time.monotonic() - start    # safe

What goes wrong

Last write wins loses data

Many replicated databases resolve conflicting writes by keeping the one with the latest timestamp. Suppose node A's clock is 100 ms ahead of node B's.

  1. A client writes x = 1 via node A, stamped 10:00:00.100.
  2. Later, in real time, another client writes x = 2 via node B, stamped 10:00:00.050.
  3. The system compares timestamps and keeps x = 1.

The newer write is silently discarded, with no error. See eventual consistency explained.

Leases and locks expire at the wrong time

A node holds a lock "for 10 seconds". If its clock runs slow, or the process pauses, it may believe it still holds the lock after everyone else considers it expired. See how distributed locks work.

Other casualties

  • Logs from different machines appear out of order, making incidents hard to reconstruct.
  • Tokens and certificates are rejected as "not yet valid" or "expired" when clocks disagree.
  • Caches expire too early or too late.
  • Time-ordered IDs can collide or go backwards after a clock step.

Logical clocks: order without time

In 1978, Leslie Lamport observed that what systems usually need is not the time, but the order of events, and specifically which events could have influenced which.

He defined happened-before:

  • Within one process, events are ordered by when they occur.
  • Sending a message happens before receiving it.
  • If A happened before B and B before C, then A happened before C.

If neither event happened before the other, they are concurrent: neither could have affected the other.

Lamport timestamps

A Lamport clock is a simple counter kept by each process:

  1. Before each local event, add 1 to the counter.
  2. When sending a message, include the counter.
  3. When receiving a message, set the counter to max(local, received) + 1.

This guarantees that if A happened before B, A's timestamp is smaller. Breaking ties by process ID gives a total order everyone agrees on.

The limitation: the reverse does not hold. A smaller timestamp does not prove an event happened first, so Lamport clocks cannot detect concurrency.

Vector clocks

A vector clock fixes that. Each process keeps a counter for every process.

  • Increment your own entry on each event.
  • Send the whole vector with each message.
  • On receiving, take the element-wise maximum, then increment your own entry.

Comparing two vectors:

  • If every entry of A is less than or equal to B's, A happened before B.
  • If each is larger in some entry, the events are concurrent.

That is exactly what a replicated database needs to know: is one version newer, or were they written independently and so in conflict? Dynamo-style systems use this idea (as version vectors) to detect conflicting writes instead of guessing with timestamps.

Wall-clock timestampLamport timestampVector clock
SizeOne numberOne numberOne number per node
Reflects real timeApproximatelyNoNo
Orders causally related events correctlyNot reliablyYesYes
Detects concurrent eventsNoNoYes

Hybrid logical clocks

A hybrid logical clock combines physical time with a logical counter. Timestamps stay close to real time, which is useful for people and for snapshot reads, while still respecting cause and effect. CockroachDB, YugabyteDB and MongoDB use variants.

TrueTime: admitting uncertainty

Google's Spanner database takes a different route, described in the Spanner paper. Its data centres have GPS receivers and atomic clocks, and its TrueTime API does not return a single time. It returns an interval: "the true time is definitely between these two values".

The interval is kept small, a few milliseconds. When a transaction commits, Spanner picks a timestamp and then waits until that timestamp is definitely in the past everywhere before reporting success. That brief wait guarantees that transaction timestamps reflect real-time order across the whole planet, giving strong consistency at global scale.

The lesson generalises: a clock reading is a measurement with an error bar. Safe systems account for the error instead of pretending it is zero.

Practical advice

  1. Run time synchronisation everywhere and monitor clock offset. Alert when it grows.
  2. Use monotonic clocks for timeouts, durations and rate limiting.
  3. Do not order events from different machines by wall-clock timestamps when correctness depends on it.
  4. Prefer versions and sequence numbers issued by a single authority, such as a database's log position.
  5. Allow for skew when checking token and certificate validity.
  6. Store timestamps in UTC. Time zones are a separate source of bugs; see why time zones are a programmer's nightmare.

Frequently asked questions

What is clock drift?

The gradual gain or loss of time by a computer's clock compared with true time, caused by imperfections in its oscillator.

How accurate is NTP?

Typically within a few milliseconds on a good local network and within tens of milliseconds over the internet. It can be much worse under poor conditions.

What is the difference between a Lamport clock and a vector clock?

A Lamport clock is a single counter that orders causally related events. A vector clock keeps a counter per node and can also tell when two events are concurrent.

Why not just use very accurate clocks?

They help, and some systems do, but any clock has some error. The system must still account for that uncertainty, as Spanner does by waiting it out.

Conclusion

In a distributed system, "what time is it?" has no exact answer, and "which came first?" cannot be settled by comparing timestamps from different machines. Use physical clocks for what they are good at, showing times and measuring local durations, and use logical clocks, versions or bounded-uncertainty time whenever correctness depends on order.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy