In April 2016 an engineer filed CASSANDRA-11586, "Avoid Silent Insert or Update Failure In Clusters With Time Skew": when the coordinator's clock lags, an UPDATE stamped older than the existing value is thrown away and the client is told it succeeded. The ticket is still open, and no clock in it is broken. Every clock API involved did exactly what it promises, and none of them promises order.

That is the argument of this note. A created_at column looks like a total order because it sorts, but the wall clock behind it was never built to order events. The fix is a primitive whose contract actually says "order" (a logical clock, a fencing token), not a tighter NTP configuration.

Most engineers read a timestamp as a sequence number

The working model goes like this: every server syncs to NTP, NTP is accurate to a few milliseconds, and writes that far apart are rare. So ORDER BY created_at is order, give or take a coin flip on near-simultaneous events that nobody cares about.

That model is useful for logs and dashboards. It fails in three places the reader rarely checks. A single machine has two clocks with different contracts. Machines drift apart by amounts that depend on hardware and hypervisor more than on NTP. And the worst errors come from clocks that step or freeze, not from ones that drift slowly.

Clocks drift, step and freeze, even on one machine

Two clocks, two contracts

Linux exposes both, and the clock_gettime(2) man page is precise about the difference. CLOCK_REALTIME is a "settable system-wide clock" that can jump in either direction when an administrator or a daemon sets it. CLOCK_MONOTONIC cannot be set, never jumps, and is only nudged in frequency; it also stops counting while the machine is suspended. Both are documented as non-decreasing, but only the monotonic guarantee survives someone setting the time.

Every mainstream runtime maps onto that split, though not always visibly:

RuntimeWall clock (can jump)Monotonic (elapsed time only)
LinuxCLOCK_REALTIMECLOCK_MONOTONIC
Go 1.9+wall reading inside time.Now()monotonic reading inside time.Now(), used by Sub and comparisons
Java 21System.currentTimeMillis()System.nanoTime(), "not related to any other notion of system or wall-clock time"
PostgreSQL 18clock_timestamp()none exposed

The Go row has a history. Before Go 1.9, time.Now() carried only the wall reading. On 1 January 2017, Cloudflare's RRDNS computed a round-trip time as the difference of two wall-clock reads, the leap second made that difference negative, and the negative value reached rand.Int63n(), which panicked. About 0.2% of DNS queries failed until a fix rolled out. The code had assumed that wall time only moves forward. Go's time package now embeds a monotonic reading in every time.Now() precisely so that this assumption holds for subtraction.

Postgres stamps the start, not the commit

The PostgreSQL row hides a quieter trap. now(), CURRENT_TIMESTAMP and transaction_timestamp() all return the time the transaction started, and the documentation calls this "a feature". Only clock_timestamp() advances during a transaction.

Follow that through for a column declared created_at DEFAULT now(). Transaction A begins at 10:00:00.100 and commits at 10:00:00.900. Transaction B begins at 10:00:00.300 and commits at 10:00:00.400. A reader polling at 10:00:00.500 with WHERE created_at > $last_seen sees B, advances its cursor past 10:00:00.300, and never sees A. One machine, one clock, no skew at all, and the ordering is still wrong, because the timestamp records when work started while visibility follows commit.

The hardware drifts and the hypervisor freezes

A quartz oscillator drifts in proportion to its frequency error, and 1 ppm is 86.4 ms a day. The figures below are manufacturer spec ranges, not measurements.

Standard crystal at ±20 ppm, unsynchronised
1.73s/day

Upper end of the typical ±10–20 ppm spec for commodity oscillators.

Oven-controlled crystal at ±0.05 ppm
4.3ms/day

The hardware you buy when drift has to be bounded without a network.

NTP exists to cancel that drift, and RFC 5905 specifies how. Small corrections are slewed by adjusting frequency; large ones are stepped with settimeofday(). A step is a jump in CLOCK_REALTIME, and every timestamp taken just after it can be earlier than one taken just before.

Virtualisation adds a third behaviour: freezing. VMware's own KB 335071 says a vMotion or snapshot resume can leave a guest "tens or even hundreds of milliseconds" behind. VMware Tools steps the clock only past 1000 ms; below that, NTP slews it back, and a 100 ms error can take around 200 seconds to close. For those minutes, that VM stamps every row in the past.

And sometimes nothing corrects it. A practitioner postmortem from September 2026 describes chrony silently failing to start after an Azure-to-AWS migration. The clock drifted for 67 days until it crossed the five-minute validity window of AWS SigV4 signatures, and every AWS API call failed at once. Drift gives no warning until something downstream has a hard threshold.

Across machines, skew is hundreds of microseconds at best

The honest question is how far apart well-run clocks actually sit. Vendors publish bounds; independent measurements are rarer and do not always agree with them.

Measured clock offset by daemon and platformIndependent measurements, different methods. Facebook: production fleet against a GNSS and atomic-clock reference, 1PPS phase comparison (2020). AWS and Azure chrony: chrony tracking logs over 12–24 h, a few instances, one Azure region, 2021–2023 (Liberty Systems). Azure guest PTP: typical maximum error, small VMs in six regions against NTP/GPS-referenced clients, March 2025 (Sommars). Small samples; read as orders of magnitude, not a leaderboard. · lower is better
chrony, AWS Fargatebest case
70µs
chrony, Facebook fleet100 with HW timestamps
200µs
chrony, Azure VM15-min spikes
200µs
chrony, AWS EC2median
250µs
ntpd, Facebook fleetlab: up to 13 000
1 500µs
Azure guest PTPone VM over 70 000
20 000µs

Two readings of that chart matter. First, even the good bars are hundreds of microseconds. Two writes on different nodes that land inside that window have no trustworthy order, however well the fleet is run. Facebook's move from ntpd to chrony bought a 10x improvement and did not change that fact.

Second, the last bar is a vendor claim that failed in the field. Azure markets its guest PTP at low-microsecond accuracy; Steven Sommars measured 5–20 ms of typical error, one to four orders of magnitude off, and reported it to Microsoft in February 2025. AWS, by contrast, states an error bound under 100 µs over NTP and under 40 µs with its PTP hardware clock, without disclosing a method. A single-instance spot check by Yugabyte, a vendor with a stake in clocks, saw about 27 ns with the hardware clock. That is directional, not a benchmark. The pattern is that a datasheet tells you what is possible, and only your own measurement tells you what you have.

Neither result rescues created_at. A 27 ns clock still steps on a leap second, still reads whatever the hypervisor tells it, and still stamps the transaction start in Postgres.

Last-write-wins turns skew into silent data loss

Skew on its own only makes timestamps inaccurate. It becomes data loss when a system uses the timestamp to decide which write survives.

Cassandra resolves conflicting writes by keeping the one with the highest timestamp. Take two clients updating the same row a few milliseconds apart through different coordinators, where the second coordinator's clock is 50 ms behind. The later update arrives with the earlier timestamp and loses. That is exactly the failure in CASSANDRA-11586, and the defining detail is the response: success. There is no error to retry and no log line to investigate. Kyle Kingsbury's 2013 teardown, "The Trouble with Timestamps", traced the same class of loss through earlier Cassandra tickets and a reverted patch from 2010.

Skew can also spread. CASSANDRA-11991 documents a node whose bad clock leaked into other nodes through the lightweight-transaction path, a side effect of an earlier fix. A timestamp-based system has no way to tell a fast clock from a genuinely later write, so it trusts the fast clock.

Dynamo took the opposite route. When its version metadata shows two writes are concurrent, it keeps both as "siblings" and hands them to the client to merge, the classic example being a shopping cart unioned on read. That costs application code. It never discards a write the client was told had succeeded.

Causality orders events without asking a clock

Leslie Lamport's 1978 paper removed physical time from the question. Event a happens before b if they are in program order on one process, or if a sends a message that b receives, closed transitively. Anything else is concurrent, and the resulting order is partial, which is the honest answer: some events have no order.

The clock that respects it fits in a few lines:

lamport.ts
let counter = 0;
 
// Before every local event, and before attaching the value to an outgoing message
export const tick = () => ++counter;
 
export function receive(remote: number): number {
  counter = Math.max(counter, remote) + 1;
  return counter;
}

If a happened before b, then a's counter is smaller. The converse does not hold: two unrelated events get different numbers anyway, so a Lamport order is total but partly arbitrary. That is the property created_at pretends to have and does not. Lamport numbers at least never put an effect before its cause.

Vector clocks keep one counter per node and merge by taking the elementwise maximum. That buys the ability to say "these two are concurrent", which is what Dynamo and Riak use to produce siblings. It also carries the failure that shaped everything after. Riak originally used client IDs as vector entries, so a vector grew with every client that ever wrote the object. Riak's own retrospective lists the results: sibling explosion, slow reads, nodes crashing out of memory. Riak 1.0 (September 2011) switched to per-vnode actors, bounding a vector to roughly the replica count, and later added Dotted Version Vectors to prune stale siblings.

Hybrid Logical Clocks are the compromise the field settled on. An HLC value stays close to the local wall clock but carries a logical counter that advances whenever physical time has not moved forward or an incoming timestamp is ahead. The result sorts roughly like wall time and never violates causality. CockroachDB keeps an HLC on every node, "always greater than or equal to wall time", folds in the timestamp of every incoming request, and assigns transaction timestamps from it. MongoDB's SIGMOD 2019 paper explains shipping one in 3.6 by elimination: pure Lamport clocks lose any link to physical time, and vector clocks cost O(N) per message.

Spanner shows what the other path costs. TrueTime returns an interval rather than an instant, and a commit waits until the interval has safely passed. The OSDI 2012 paper reports that uncertainty as a sawtooth of roughly 1–7 ms, driven by polling time masters every 30 seconds under an assumed worst-case drift of 200 µs per second. That assumption is about ten times the spec of a cheap crystal. Google engineered for the pathological clock, and it needed GPS receivers and atomic clocks in every datacenter to make waiting affordable.

A lease is a bet on clocks; a fencing token cancels it

Locks are where the clock problem stops being about sorting and starts corrupting shared state. A distributed lock almost always has an expiry, because, as Martin Kleppmann puts it, otherwise "a crashed client could end up holding a lock forever". An expiring lock is a lease, and a lease is safe only if the holder's sense of elapsed time matches the lock service's.

It often does not. A garbage-collection pause longer than the lease leaves a client believing it holds a lock that has already been granted to someone else, and both write. Kleppmann cites this as a real HBase incident. No timeout value fixes it, because the pause can always be longer than the timeout.

The fix moves the check to the resource being protected. Each grant of the lock carries a number that increases with every acquisition, and storage rejects any write whose number is lower than one it has already accepted:

fenced-write.sql (any PostgreSQL version)
UPDATE documents
SET body = $1, fence = $2
WHERE id = $3 AND fence < $2;
-- 0 rows: a later lock holder has already written; abort

The paused client wakes up with token 33, storage has already seen 34, and the write bounces. Chubby shipped this in 2006 as a "sequencer"; ZooKeeper's zxid or a znode version serves the same role. etcd's lock guide does not use the term, but a key's creation revision is monotonic and is used as a token in practice.

Redis's Redlock has no such number, which is the core of Kleppmann's objection: its safety rests on bounded clock error and bounded pauses, and neither is guaranteed. antirez replied that Redlock only needs to count an interval to within about two seconds and rechecks elapsed time after acquiring a majority. He conceded one point, that implementations should use a monotonic clock, and disputed the rest. Neither side has retracted since. Kleppmann's own split is the useful part: a lock for efficiency can tolerate an occasional double execution, and a lock for correctness needs a token that storage enforces.

Where this explanation stops

HLCs do not remove physical time; they lean on it. The wall-clock component is still read from a synchronised clock, so an HLC-based system still needs skew kept within some bound, and a node far outside it degrades the "close to wall time" half of the promise.

Causal order is also narrower than real-time order. If a user writes on one continent and a friend reads on another with no message between them, causality says nothing about which came first. Spanner's external consistency covers that case; most applications do not need it, and the ones that do pay in hardware or latency.

Tight synchronisation is no longer exotic either. Facebook's chrony numbers, AWS's hardware clocks and Meta's PTP work mean "atomic clocks or nothing" is a cost argument now, not a physical one. Better sync narrows the window where timestamps mislead, and it helps with logs and TTLs. It does not turn a timestamp into a sequence.

The discontinuities are also changing shape. The BIPM voted in November 2022 to stop leap seconds by 2035, but Duncan Agnew's 2024 Nature paper projects that Earth's rotation would first require a negative leap second, around 2029 after polar ice melt pushed it back from about 2026. No production system has ever lived through one.

What follows for your schema and your locks

  • Treat created_at as a label, not a cursor. Anything that pages, syncs or replicates "everything since X" by timestamp can skip rows, even on one Postgres primary, because now() records transaction start. Order by something the writer assigns in commit order, or make readers tolerate late arrivals by overlapping their window.
  • Never let a wall-clock comparison pick a winner. If your store resolves conflicts by last-write-wins, assume some acknowledged writes are gone and design around it: keep siblings, use version checks, or accept the loss knowingly.
  • Measure durations with the monotonic clock. Timeouts, retries and latency metrics built on wall time will eventually go negative or jump; Cloudflare's 2017 outage is the reference case.
  • Put a fencing check next to every lease that guards correctness. The lock service hands out the number, and the storage layer refuses anything stale. Without the second half, the first is decoration.