Data engineering
Part 5 of 5 · CDC & DebeziumCDC Failure Modes — Lag, Schema Breaks, Tombstones, Backfills & Replays
CDC fails in production as slot disk, breaking DDL, missing tombstones, and a panicked offset rewind. Freshness is an SLO. Schema changes expand then contract. Rebuild a projection beside the live one, then swap.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Lag is allowed to be visible
Prefer
Backpressure, then a sized sink
A slow sink should lag. Dropping events to look fresh is data loss. Scale, chunk the snapshot, or shed a non-critical table on purpose.
- Freshness SLO is event time to apply time.
- Slot retained bytes page before the disk fills.
- Rebuild into a shadow, then swap.
Alternative
Edit the offset under stress
Skipping LSN hides the incident and opens a gap. Rewinding without an idempotent sink double-applies. Both fail the postmortem.
- Manual offset edits are not a mitigation.
- In-place poison of the live index is hard to undo.
- Schema history is not disposable.
Overview
Happy-path Debezium demos are common. Production is a slot disk alarm, a DDL change that wedges the connector, a search index that still shows deleted users, and a rewind that double-charges because the sink was not idempotent. This lesson is the runbook for the cluster.
You already know capture is at-least-once. You already know an unconsumed slot retains WAL. Here you practice the response.
Failure catalog
| Failure | Symptom | Why | Mitigation |
|---|---|---|---|
| Connector lag | Sink minutes or hours behind | Slow sink, too few tasks, snapshot locks | Lag SLO, scale the sink, chunk snapshots |
| Slot WAL growth | Disk alarm on the primary | Consumer stopped, no heartbeats, huge backlog | Alert on retained bytes, page, heartbeats |
| Schema break | Connector crash or poison events | Incompatible DDL, history gap | Expand/contract, history retained |
| Missing deletes | Row returns in search | No tombstone, or only a soft delete with no convention | Emit deletes; tombstones on compacted topics |
| Bad replay | Duplicates or gaps | Offset rewind, sink not idempotent | LSN guard or inbox; staged rebuild |
| PII leak | Wide table capture | Include list too broad | Mask or drop columns; minimize capture |
| Ordering surprise | Cross-table rule broken | Many tables, many partitions | Outbox domain event; one aggregate key |
Lag SLOs
Pick a freshness target by tier. Example: p99 of event time to sink apply under 30 seconds for tier-1 tables. “The task is RUNNING” is not that metric.
Measure source lag (LSN or time), Connect task lag, sink write latency, and consumer lag. If the sink is slow, lagging is the correct behavior. Dropping is not. Scale the sink, batch the bulk API, or intentionally remove a non-critical table from the publication.
Decisions
- 1
1. Lag SLO breach
- next2. Where is time spent
- ?
2. Where is time spent
- slot3. WAL rate locks heartbeats
- tasks4. Task errors and restarts
- sink5. Batch or scale the sink
- 3
3. WAL rate locks heartbeats
- next6. Page and mitigate
- 4
4. Task errors and restarts
- next6. Page and mitigate
- 5
5. Batch or scale the sink
- next6. Page and mitigate
- 6
6. Page and mitigate
- next7. Capacity bug or schema
- 7
7. Capacity bug or schema
Lesson map
CDC Failure Modes — Lag, Schema Breaks, Tombstones, Backfills & Replays
CDC fails in production as slot disk, breaking DDL, missing tombstones, and a panicked offset rewind. Freshness is an SLO. Schema changes expand then contract. Rebuild a projection beside the live one, then swap.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Lag SLO breach"] b["2. Where is time spent"] c["3. WAL rate locks heartbeats"] d["4. Task errors and restarts"] a -->|1. Lag SLO breach| b b -->|slot| c b -->|tasks| d
Schema breaks
- Expand. Add a nullable column or a new field. Writers start populating it. Consumers ignore fields they do not know.
- Migrate. Backfill. Dual-read if a reader still needs the old shape.
- Contract. Remove the old field only after every consumer and the connector history can tolerate the absence.
Breaking moves: in-place type changes, renames without an alias, dropping a column a consumer still reads, recreating a table so OIDs change under a connector that keyed on them.
If the schema history topic is lost, the connector often cannot decode. Recovery is a snapshot rebuild, not a hope. History is ops state. How a schema registry scores backward versus forward compatibility is Schema Evolution, Compatibility & Dead Letter Queues. Follow that page for the registry. Follow this page for connector history and deploy order.
Tombstones and deletes
Flow
- 1
1. DELETE id 42 in the database
- next2. Debezium reads the delete
- 2
2. Debezium reads the delete
- next3. Key 42 with a delete envelope
- 3
3. Key 42 with a delete envelope
- next4. Tombstone key 42 null value
- 4
4. Tombstone key 42 null value
- next5. Compaction drops older values
- 5
5. Compaction drops older values
- next6. Sink removes document 42
- 6
6. Sink removes document 42
Without the tombstone, compaction can leave an older value for that key, and a rebuild resurrects the row. Without a delete apply, the sink never removes the document even if the topic is correct.
Soft deletes need an explicit convention (op, or a deleted flag) that every sink implements. “We only soft-delete” is not a tombstone story until the sink treats the flag as a removal.
Backfills and replays
| Mode | Use | Risk |
|---|---|---|
| Incremental or ad hoc snapshot | Add a table, repair a subset | Load on the primary |
| Full connector snapshot | New pipeline or total rebuild | Long IO window |
| Offset rewind | Reprocess from an LSN or time | Duplicates; sink must be idempotent |
| Rebuild from the compacted topic | Projection rebuild | Needs retained history or a new snapshot |
| Dump plus CDC catch-up | Analytics | Cutover watermark discipline |
Rule: do not rewind offsets into a non-idempotent sink. Prefer a new index and an atomic alias swap, or a shadow table and a cutover. The idempotent apply you are counting on is Exactly-Once CDC Pipelines. If the rows are domain events, identity comes from Transactional Outbox & Inbox Patterns.
Ops checklist
- Alert on replication slot retained WAL bytes and age.
- Alert on connector failed and on task restart loops.
- Alert on sink lag versus the SLO, by table tier.
- Runbook: drop non-critical tables from capture during an incident, on purpose.
- Runbook: schema change paired with consumer deploy order (expand before writers, contract after readers).
- Runbook: shadow rebuild and swap. Do not poison the live projection in place.
- Game day: stop the connector for 30 minutes. Prove disk headroom and catch-up.
PII response: stop the connector, tighten the include list or mask columns, scrub topics if retention allows, rotate the sinks that already stored the columns.
A cross-table invariant that flickered is often the wrong capture shape. Accept the window, or publish one outbox event per business fact.
Lag budget and a tolerant reader (run this)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The reader depends on status and ignores status_reason. That is the expand step. Removing status before readers move is the contract break.
Interview Q&A
What happens if the CDC consumer stops?
Answer
On Postgres the slot retains WAL until the LSN is confirmed. Disk can fill. The primary’s availability is now coupled to a downstream consumer. Alert on retained bytes before the volume is full.
How do you change a column type safely with CDC?
Answer
Do not change it in place. Expand a new field, backfill, move readers, then contract. Confirm the connector’s schema history still decodes both shapes. Registry rules are the schema-evolution lesson.
Why tombstones?
Answer
Compaction removes older records for a key only if a tombstone (null value) tells it the key is gone. Sinks also need a delete they will apply. Missing either one brings rows back or leaves them forever.
Is an offset rewind safe?
Answer
Only when every sink is idempotent or you are rebuilding into a new target. A rewind into a live charger or mailer is a duplicate incident. Prefer shadow rebuild and swap.
Incremental snapshot versus full snapshot?
Answer
Incremental or signal snapshots repair a subset with less disruption. A full snapshot fits a new pipeline or a total rebuild. Both load the primary. Schedule them.
How do you prove freshness?
Answer
Event-time to apply-time, with an SLO per tier. Task status is necessary and not sufficient.
A cross-table invariant broke for a few seconds. Why?
Answer
Table CDC publishes each row on its own key and partition. There is no single commit visible to the consumer across tables. Use an outbox domain event when the invariant is the product contract.
PII showed up on a topic. What now?
Answer
Stop the connector, narrow the include list or mask columns, scrub retained topic data if the retention window allows, and rotate sinks. Then add the column review to the change checklist.
Someone wants to delete the schema history topic. What do you say?
Answer
No. The connector uses it to decode the table across DDL. Losing it is a rebuild, not a cleanup. Take a retention policy instead.
Pitfalls
Slot disk, a type change that crashed the connector, search showing deleted orders, and a rewind that double-applied. For each, write the first metric, the wrong fix, and the right fix. The wrong fixes are delete-the-history, skip-the-LSN, ignore-deletes, and rewind-anyway.