Data engineering
Part 1 of 5 · CDC & DebeziumChange Data Capture — WAL Tailing, Debezium & Event Pipelines
Dual-write splits one business fact across a database commit and a later publish. CDC reads the database change log so the commit is the event. This hub maps log versus poll, Debezium, the existing outbox lesson, exactly-once effects, and failure modes.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Commit is the event
Prefer
One database commit, then a log reader
The application writes once. Logical decoding or the binlog turns that commit into an event. If the transaction rolls back, there is no event.
- No second commit against the bus inside the request.
- Deletes show up as delete records, not as missing rows.
- Downstream lag is an SLO, not a silent gap.
Alternative
Update Postgres, then publish
Two systems, two failure domains. One success and one failure leaves search, warehouse, or a sibling service wrong until someone reconciles.
- Crash after COMMIT and before publish: the row exists, the event does not.
- Publish then roll back: a phantom event with no durable row.
- Retries without a stable key double-apply.
Mental model — commit equals event
The bus is transport. The database commit is the source of truth for the change.
- 1
App commits a row change
One transaction. No produce call inside it. - 2
WAL or binlog records it
The recovery log is the change feed. Crash-recovery internals stay on the storage-engine lessons. - 3
CDC connector decodes
Logical decoding or row-format binlog yields insert, update, delete, or snapshot-read. - 4
Event bus topic
Retention and fan-out. Partition keys and consumer groups live in the Kafka messaging cluster. - 5
Downstream apply
Search, cache, warehouse, or another service. Idempotent apply plus a lag SLO.
Overview
Interviewers ask how Postgres stays in sync with Elasticsearch, a cache, or another service. The weak answer is “the app writes the row and then publishes.” That is dual-write. The database and the bus do not share a commit. Partial failure is permanent drift unless you have a reconciliation job you can actually trust.
Change data capture reads the database’s own change log. The commit is the event. Rollback produces nothing. That is the whole idea of this cluster.
This page maps the decisions. It does not re-teach Kafka topics, partitions, or consumer groups. Those live in Apache Kafka — Topics, Partitions, Brokers & Consumer Groups and Delivery Semantics — At-Least-Once, At-Most-Once & Exactly-Once. It also does not re-teach crash recovery. Checkpoints, REDO, and torn pages stay in Write-Ahead Log — Durability, Checkpoints & Crash Recovery and Database Storage Engines — WAL, B-Trees & LSM Trees. Here the log is only a CDC source: logical decoding, publications, and replication slots.
Approximate aggregations are a different problem. Quantile sketches live in Approximate Aggregations — Sketches for Quantiles, Cardinality & Merge Pipelines. CDC is about whether a change arrived, not about approximating it.
What this cluster covers
- WAL Tailing vs Query-Based CDC — Log vs Poll Tradeoffs — deletes, ordering, primary load, slot retention.
- Debezium & Kafka Connect — Snapshots, Offsets, Schema History & Heartbeats — snapshot then stream, resume tokens, DDL history.
- Transactional Outbox & Inbox Patterns — domain events co-committed with the business write. Idempotency Part 6, not a page in this series.
- Exactly-Once CDC Pipelines — Idempotent Consumers, Keys & At-Least-Once Reality — keys, tombstones, inbox, LSN guards.
- CDC Failure Modes — Lag, Schema Breaks, Tombstones, Backfills & Replays — slot disk, expand/contract, shadow rebuilds.
Flow
- 1
1. App commits the row
- next2. WAL or binlog records it
- 2
2. WAL or binlog records it
- next3. Connector decodes the change
- 3
3. Connector decodes the change
- next4. Event lands on the bus
- 4
4. Event lands on the bus
- next5. Search cache warehouse service
- 5
5. Search cache warehouse service
- next6. Idempotent apply and lag SLO
- 6
6. Idempotent apply and lag SLO
Lesson map
Change Data Capture — WAL Tailing, Debezium & Event Pipelines
Dual-write splits one business fact across a database commit and a later publish. CDC reads the database change log so the commit is the event. This hub maps log versus poll, Debezium, the existing outbox lesson, exactly-once effects, and failure modes.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. App commits the row"] b["2. WAL or binlog records it"] c["3. Connector decodes the change"] d["4. Event lands on the bus"] a -->|1. App commits the row| b b -->|2. WAL or binlog records it| c c -->|3. Connector decodes the change| d
When CDC beats dual-write
| Approach | Strength | Cost | Use when |
|---|---|---|---|
| Dual-write (DB then bus) | Easy to sketch | Not atomic; silent drift | Never on a critical path |
| Query-based CDC (poll) | Works on almost any DB | Misses hard deletes; load; jitter | No log access, or a tiny soft-delete table |
| Log-based CDC | Ordered, delete-aware, light on queries | Slot and retention ops; schema evolution | Default OLTP to events |
| Outbox plus CDC | Same transaction as the business write | Extra table and a relay | The event is not the raw row |
| 2PC / XA | Theoretical atomicity | Availability and coordinator tax | Almost never at web scale |
Rule: prefer log-based CDC when the downstream system should follow table state. Prefer the transactional outbox when the business event is not one-to-one with a row mutation, or when you need a stable public schema the OLTP table does not have.
Decision path
Decisions
- 1
1. Downstream must follow DB state
- next2. Can we read the log
- ?
2. Can we read the log
- yes3. Log-based CDC
- no4. Query poll and accept limits
- 3
3. Log-based CDC
- next5. Event shape equals the row
- 4
4. Query poll and accept limits
- ?
5. Event shape equals the row
- yes6. Debezium table capture
- no7. Outbox row plus relay
- 6
6. Debezium table capture
- next8. Idempotent consumer design
- 7
7. Outbox row plus relay
- next8. Idempotent consumer design
- 8
8. Idempotent consumer design
- next9. Lag schema tombstone runbooks
- 9
9. Lag schema tombstone runbooks
- Yes, we can read the log, and the event is the row. Table capture. Search indexes, caches, and warehouse mirrors.
- Yes, we can read the log, and the event is a domain fact. Insert an outbox row in the same transaction. Relay it. Do not also publish from the request thread.
- No log access. Poll a watermark. State the delete hole and the primary load out loud.
- 2PC. Name it, then decline it. Coordinator failure and heuristic outcomes cost more than an outbox backlog.
What “the commit is the event” rules out
- Phantom events from a publish that happened before
ROLLBACK. - Lost events from a commit that happened before a failed publish, if you are tailing the log. Dual-write still loses them.
- Treating CDC as free. You still owe lag SLOs, a replay story, a schema contract, and a tombstone story. Those are the last two lessons.
Capturing every table “just in case” widens PII blast radius, cost, and noise. Start from the downstream contract, then the include list.
In-memory dual-write versus CDC (run this)
No database and no bus. The first path commits the row and then fails the publish. The second path emits only when the commit flag is true.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The token is a teaching stand-in. Production sinks usually dedupe on an outbox event_id or on (primary key, LSN). That design is Exactly-Once CDC Pipelines.
Interview Q&A
What is change data capture?
Answer
Propagating database changes by reading the database change stream, not by dual-writing from the application. The unit is a committed row change or a logical event derived from that commit.
Why is dual-write dangerous?
Answer
Two independent systems and no shared commit. If the database succeeds and the publish fails, downstream never hears. If the publish succeeds and the database rolls back, downstream acts on a row that does not exist. You will not notice until reconciliation.
Log-based versus query-based?
Answer
The log is commit-ordered, delete-aware, and usually lighter on the primary. A poll is portable, misses hard deletes, and spends lag on the poll interval plus SELECT load. Depth is the next lesson.
Where does Kafka fit?
Answer
Transport and retention for change events. It is not a substitute for capturing the commit. Consumer groups, partitions, and delivery knobs stay in the messaging cluster. This cluster applies them to CDC keys, offsets, and tombstones.
Outbox versus plain table CDC?
Answer
Table CDC is the right contract when the projection should mirror the row. The outbox is the right contract when the public event is a domain fact, possibly spanning several tables, with its own schema. See Transactional Outbox & Inbox Patterns.
Is CDC exactly-once?
Answer
Capture through the connector is usually at-least-once, because offsets flush in batches and restarts redeliver. Exactly-once effects need idempotent sinks and stable keys.
What breaks first in production?
Answer
Replication slot disk growth, schema changes without a compatibility plan, lag with no SLO, and deletes that never become tombstones. The failure-modes lesson is the runbook.
Why not two-phase commit between Postgres and Kafka?
Answer
XA gives a theoretical atomic boundary and an operational one you rarely want: coordinator failure, long locks, heuristic outcomes. The pragmatic boundary is one database commit plus an idempotent consumer.
What is logical decoding, in one line?
Answer
A Postgres facility that turns WAL into a row-level change stream for a replication slot. It is a CDC prerequisite, not the crash-recovery algorithm. Recovery stays on the WAL durability lesson.
Should you capture every table?
Answer
No. Every included table is cost, PII, and a schema you must evolve. Start from the downstream contract.
Pitfalls
On one whiteboard column, write dual-write with an X between COMMIT and publish. On the other, write a single COMMIT and a log reader. For each column, say what exists in the table and what exists on the bus. Then mark whether the downstream event should be the row or a domain fact, and name the lesson that owns that choice.