Ask which outbox relay to use and the usual answer ranks them by latency: polling is "seconds", change data capture is "milliseconds". Nobody has measured that with a stated method, and every figure I could trace is either an assertion or poll-interval arithmetic. What the sources do document is how each relay fails without telling you. A naive polling relay can skip committed rows forever, and a replication slot on an idle RDS instance retained about 18 GB of WAL a day until the disk filled.
Both relays give you the same guarantee. The choice between them is decided by which operational bill you can pay and watch: polling pays with your database's own load, bloat and an ordering trap you must code around, while log-based CDC pays with WAL the database is not allowed to delete and a failover story you have to verify yourself.
Two writes to two systems have no safe order
This is the code most services start with, here with pg 8 and kafkajs 2.2:
await db.query('BEGIN');
await db.query('INSERT INTO orders (id, total) VALUES ($1, $2)', [id, total]);
await db.query('COMMIT');
// (A) the process dies here: the order exists, the event never leaves
await producer.send({
topic: 'orders',
messages: [{ key: id, value: JSON.stringify({ id, total }) }],
});
// (C) send() times out: was the message written or not?If the process dies at (A), the state changed and nobody downstream hears about it. Swap the blocks and you get case B: the send succeeds, the commit fails, and consumers act on an order that does not exist. Confluent's write-up of the dual-write problem puts it plainly: "Reordering the operations still results in inconsistency; it's just different."
Case C is the one Node code hides. KafkaJS times a request out after 30 seconds by default, and a timeout does not say whether the broker kept the message. Retry and you may publish twice; give up and you may have published nothing. Kafka's idempotent producer, on by default in the Java client since Kafka 3.0 and still flagged experimental in KafkaJS, dedupes the producer's own retries to the broker. It knows nothing about the database commit that came before.
Concurrency adds a failure with no error at all. In Kleppmann's 2015 example, two clients write the same key to two stores; the database applies B last while the other store applies A last, and the two "will permanently remain inconsistent". No retry policy repairs that, because nothing failed.
The outbox buys atomicity and nothing more
The outbox moves the event into the same transaction as the state change:
await db.query('BEGIN');
await db.query('INSERT INTO orders (id, total) VALUES ($1, $2)', [id, total]);
await db.query(
`INSERT INTO outbox (aggregateid, type, payload) VALUES ($1, 'OrderCreated', $2)`,
[id, { id, total }],
);
await db.query('COMMIT');Now the event exists exactly when the order does, or in microservices.io's wording, it is "sent if and only if the database transaction commits". Something still has to carry rows from Postgres to the broker, and that relay can crash after publishing but before recording that it did. Every relay is therefore at-least-once, whichever kind you pick.
That symmetry is why latency and elegance are the wrong criteria. The two relays differ in where they read. A poller reads the outbox table with ordinary queries, so it competes with your workload under the rules of MVCC: locks, snapshots, vacuum. A log tailer reads the write-ahead log through a replication slot, and Postgres keeps every byte the slot has not confirmed. One relay spends the database's CPU and I/O; the other spends its disk and your attention.
| Polling relay | Log-based relay (CDC) | |
|---|---|---|
| Reads from | the outbox table, on an interval | WAL, through a logical replication slot |
| Order it sees | id or transaction-start order | commit order |
| Added latency | up to one poll interval, T/2 on average | no independent measurement found |
| How it fails silently | skipped rows, a stalled cursor, a growing backlog | an inactive slot retaining WAL, a gap after failover |
| Extra moving parts | a worker process | wal_level=logical, a connector, usually on the JVM |
| Outbox table housekeeping | delete or partition published rows | insert and delete in one transaction; the table stays empty |
Polling pays with your database's own resources
The obvious poller keeps a cursor and asks for id > $last. Postgres hands out the sequence value at INSERT, but the row becomes visible at COMMIT, and the sequence documentation promises nothing about lower values committing first. Transaction A takes id 100, B takes 101, B commits first. The poller reads 101, advances its cursor, and never sees 100.
The correct poller stores the writing transaction's id and reads only below the oldest snapshot still open:
SELECT * FROM outbox
WHERE transaction_id < pg_snapshot_xmin(pg_current_snapshot())
AND (transaction_id, position) > ($1, $2)
ORDER BY transaction_id, position
LIMIT 500;Every transaction id below pg_snapshot_xmin has either committed or aborted, so nothing can appear behind the cursor later. The fix costs liveness. It orders by transaction start rather than commit, and one long-running transaction pins xmin, so the relay stops delivering every newer event until that transaction ends. Sequin, a company that sold CDC, documented the stall in 2024; the mechanism follows from the query whoever reports it. The trap keeps being rediscovered: a ruby-dcb pull request dated 29 September 2026 fixes a store that "permanently skips events committed out of position order".
The table also degrades as a queue does. An independent test by msdousti loaded a one-million-row outbox and churned it from parallel sessions:
- Fresh table
- 0.133ms
- Under churn, unpartitioned
- 18 553ms
- Under churn, list-partitioned on published_at
- 2.5ms
The query did not change between the first two bars; the table's history did. The slow run made 96.6 million heap fetches against zero on the fresh table, because every processed row leaves versions behind that the query still has to visit. Partitioning brought it down to 1,000. Nobody gets that for free: a polling outbox needs a retention design from its first week.
Vacuum has its own dependency on long transactions. PlanetScale ran a generic SKIP LOCKED job queue, not an outbox, on Postgres 18 with 8 workers and overlapping 120-second analytics queries that kept the xmin horizon pinned. PlanetScale sells the throttle that fixed it, so read these as vendor numbers:
- 383 000
VACUUM reported as blocked for the whole run; PlanetScale PS-5, 15-minute runs.
- 155 0000
Before and after PlanetScale's own Traffic Control throttle.
The same long transaction now hurts the poller twice: it stalls the xmin cursor and it stops vacuum from cleaning the table the cursor reads. SKIP LOCKED does not help with either. The SELECT documentation says it "provides an inconsistent view of the data" and suits queue-like tables only for avoiding lock contention, which is a different problem from ordering.
Log-based CDC pays with WAL the database cannot release
Logical decoding emits only committed transactions, in commit order, so the ordering trap disappears. With a log reader the application can insert the outbox row and delete it in the same transaction: the WAL keeps the insert, the table stays empty, and Debezium's Outbox Event Router ignores the delete. SeatGeek went further and dropped the table entirely, writing events with pg_logical_emit_message(true, …), which puts a message into WAL as part of the transaction. They report that the router assumed an outbox table, so they wrote their own transforms.
The bill arrives through the slot. The logical decoding documentation says slots "persist across crashes and know nothing about the state of their consumer(s)". They hold back WAL and the catalog rows vacuum would otherwise remove, and "in extreme cases this could cause the database to shut down to prevent transaction ID wraparound". The two safety valves, max_slot_wal_keep_size and Postgres 18's idle_replication_slot_timeout, ship disabled: the first defaults to -1, unlimited, the second to 0, off.
An idle database is the counterintuitive case:
- 18GB/day
Gunnar Morling, November 2022; a 200 GiB disk filled in under two weeks.
The arithmetic explains it. RDS writes its own heartbeat every five minutes, archive_timeout forces a segment switch on the same cadence, and RDS uses 64 MB segments: 288 switches a day at 64 MB each is about 18 GB. With no user writes the connector has nothing to confirm, so the slot never advances. The fix is a heartbeat the connector does see, such as pg_logical_emit_message(false, 'heartbeat', now()::varchar) on Postgres 14 and later, or Debezium's heartbeat.action.query.
Failover is the second silent gap. Before Postgres 17, logical slots existed only on the primary. Infobip ran Postgres 16 with Patroni 3.3.1 and reported that after promotion "a replacement slot starts beyond Kafka Connect's durable LSN, creating a potential event gap". Their advice is to treat the slot and Debezium's stored offset "as one distributed checkpoint". Postgres 17's failover slots narrow the gap, but only slots marked synced survive promotion, and Infobip still treats the check as yours to do.
The connector is a service with its own upgrade path. Debezium 3.7, released on 29 September 2026, needs Java 17 for connectors and Java 21 for Debezium Server, which can at least run without Kafka. Zalando stayed on Debezium 2.7.4 for about two years because a keepalive change blocked upgrades, until fixes landed in 3.4.0. The lighter alternatives are less stable than they look: Sequin, a single-binary CDC tool often recommended as the anti-Debezium, was acquired in August 2025, shut its cloud in October and has had no upstream maintenance since February 2026. SeatGeek's own conclusion was that the setup took significant development and cross-team effort, and fits new projects better than retrofits.
Both relays hand the consumer duplicates
A polling relay that crashes between publishing and marking the row sends it again. A slot can return to an earlier LSN after a crash and re-send changes, which the Postgres documentation states outright. So the consumer has to dedupe, and the reliable place to do it is the consumer's own database, in the same transaction as the effect:
await db.query('BEGIN');
const { rowCount } = await db.query(
'INSERT INTO processed_events (event_id) VALUES ($1) ON CONFLICT DO NOTHING',
[eventId],
);
if (rowCount === 1) await applyOrderCreated(db, payload);
await db.query('COMMIT');Broker-side deduplication does not replace this, because every window it offers is bounded. SQS FIFO drops duplicates within five minutes of the same deduplication id, and a resend after that window goes through. BullMQ ignores a repeated jobId only while the earlier job still exists; jobs removed by removeOnComplete "will not be considered as duplicates". RabbitMQ has no deduplication and recommends idempotent consumers instead.
Ordering survives only per key. Use the aggregate id as the Kafka message key, as the Event Router does with aggregateid, and all events for one order land in one partition in relay order. That makes the relay's order the one that matters. Log tailing gives commit order from a single slot reader. Polling gives transaction-start order at best, and Piontko notes that several relay instances racing over the same table reorder events, so a polling relay should run as one active instance or split work by aggregate.
Pick the relay whose failure you can watch
No outbox: a Postgres-backed job queue such as pg-boss or Graphile Worker
Enqueue in the business transaction and there is no second system to keep consistent. This stops helping the moment the consumer is another service.
Polling outbox with an xid8 cursor and a partitioned table
No new infrastructure and no WAL risk. Budget a delay of one poll interval, an alert on backlog age, and a ceiling on transaction duration so xmin cannot pin the cursor.
Log-based CDC
These queries stall the xmin cursor and block vacuum on the outbox table at once. A log reader sees commits in order regardless of what else is open.
Debezium with the Outbox Event Router
The slot monitoring, heartbeats and upgrade discipline are already paid for, so the outbox adds a transform rather than a new system.
Polling, until failover slots are in place and rehearsed
Without synced slots a promotion can open an event gap with no error. Neon's documentation, for example, says it drops slots inactive for about 40 hours and blocks scale-to-zero while a subscriber is connected.
What would change the answer
The whole argument rests on Kafka being unable to join a transaction with Postgres. KIP-939, Kafka participation in two-phase commit, is accepted but did not ship in Kafka 4.2, and its implementation pull request was still open in April 2026. Even once it ships, Morling's counterargument holds: with an outbox "a service only needs a single resource to be available", its own database.
Two smaller events would shift the branches. If Postgres or the major hosts start shipping a finite max_slot_wal_keep_size by default, the worst CDC failure becomes a broken slot instead of a full disk, which is easier to live with. If someone publishes a polling-versus-CDC latency benchmark with stated hardware and method, latency may earn a place among the criteria; until then it has not.