Change Data Capture: Keep Two Databases in Sync
Change data capture explained: log vs triggers vs polling, ordering, duplicates, schema changes, CDC vs dual writes, and when to skip it in a live migration.

On this page · 7 sections
- What is change data capture (CDC)?
- How does CDC work: logs, triggers or polling?
- Why use CDC when migrating a live platform?
- What are the common pitfalls: ordering, duplicates and schema changes?
- CDC vs dual writes: which is safer?
- When is change data capture overkill?
- Want a second opinion on your migration?
Change data capture (CDC) is a technique that detects every insert, update and delete in a database and streams those changes to another system. The second database stays in sync with the first without anyone copying data by hand or running nightly exports.
In a live migration, CDC lets the old and new databases run side by side while users keep working. This guide covers how it works, where it breaks, and when you can skip it.
Key takeaways
- CDC copies changes, not whole tables, so the target follows the source with a small delay.
- Reading the database log is usually the most reliable method. Triggers and polling are simpler but have real costs.
- The hard parts are ordering, duplicates and schema changes, not the tooling.
- CDC is generally safer than dual writes because the source database stays the single source of truth.
- For small or low-traffic systems, a maintenance window and a plain export may be enough.
What is change data capture (CDC)?
CDC turns a database into a stream of events. Each time a row changes, a record describes what happened: the table, the primary key, the old values, the new values and the type of operation.
A consumer reads that stream and applies the changes to a target. The target might be a new database on a different engine, a search index, a data warehouse or a cache.
The key idea is that the application does not need to know. Your code writes to the database exactly as before, and the capture happens underneath. That is what makes CDC attractive when you are afraid to touch a legacy codebase.
A CDC setup usually has two phases:
- Initial snapshot: copy the existing data to the target.
- Continuous streaming: apply every change made after the snapshot started.
Getting these two phases to meet without gaps or overlaps is where most of the engineering care goes.
How does CDC work: logs, triggers or polling?
There are three common approaches. They trade simplicity against accuracy and load on the source.
| Method | How it works | Main strength | Main weakness |
|---|---|---|---|
| Log-based | Reads the database’s transaction log (for example PostgreSQL logical decoding or the MySQL binlog) | Captures every change, including deletes, with low load | Needs database configuration and permissions |
| Triggers | Database triggers write changes into an audit table | Works on almost any database | Adds overhead to every write and is easy to forget on new tables |
| Polling | A job queries for rows with a recent updated_at value | Very simple to build | Misses hard deletes and rows without a reliable timestamp |
Log-based CDC
This is what most teams mean by CDC. The database already records every committed change in its log for crash recovery and replication. A connector reads that log and publishes events. Open-source tools such as Debezium do this for several popular databases, and cloud providers offer managed versions.
Triggers and polling
Triggers are fine for a handful of tables. Polling is fine when you only care about rows that are created or updated and you trust the timestamps. In an inherited codebase, I don’t trust them by default: bulk scripts, raw SQL and old code paths can update rows without touching the timestamp column, and you usually can’t rule that out until you’ve read every write path.
Why use CDC when migrating a live platform?
A live migration has a hard constraint: the old system keeps taking writes until the day you switch. If you copy the data once, it is stale the moment the copy finishes.
CDC solves this by keeping the new database current while you build and test the new system against real data. When you are confident, you move traffic over, and the cutover can become a small step instead of a weekend of risk.
At avanzzada I led the full migration of a platform that applies to job openings automatically for its users, moving its whole architecture from a legacy stack to a new one while people kept using it every day. In that kind of work, the question is never just how to move the data. It is how to keep two systems telling the same story while real users change that story every minute.
I prefer moving piece by piece over a big-bang rewrite for a simple reason: a rewrite usually costs more and takes longer, because the old system must keep running in parallel anyway. CDC is one of the tools that makes the incremental path practical.
CDC helps in a few concrete ways:
- Realistic testing: the new system runs against fresh production-shaped data, not a stale dump.
- Gradual traffic moves: you can shift one feature or user group at a time, which fits an incremental approach over a big-bang rewrite.
- Easier rollback: while the old database stays current, you can send traffic back if something looks wrong.
What are the common pitfalls: ordering, duplicates and schema changes?
The tools are mature. Most of the trouble comes from three problems that tools cannot fully hide.
Ordering
Changes to the same row must be applied in the order they happened. If an update to a user’s email arrives before the insert that created the user, the target breaks or ends up with the wrong value.
Log-based tools generally emit changes in commit order, but if you fan events out across many parallel consumers, you can lose that order. A common rule: partition by primary key so all changes to one row go through the same consumer in sequence.
Duplicates
Most CDC pipelines are designed for at-least-once delivery. That means an event can arrive twice after a restart or a retry. Your consumer must be idempotent: applying the same change twice must give the same result as applying it once.
Upserts keyed on the primary key, plus storing the log position of the last applied change, are the usual defenses.
Schema changes
Someone adds a column, renames a field or changes a type in the source while the pipeline is running. Depending on the tool, the stream might break, drop the new field or pass something the target cannot parse.
Plan for this explicitly:
- Freeze or review schema changes on the source during the migration.
- Prefer additive changes: add new columns before using them, and remove old ones only after the cutover.
- Alert on pipeline failures and lag, so a silent stop does not turn into a surprise at cutover.
Verification
Do not assume the target matches. Compare row counts, checksums on key tables and spot checks on money-related records such as orders, payments and subscriptions. Run the comparison repeatedly, not once.
When I take over a system, my order is access first, then the build, then the money paths, then the risks. Having worked on the cart and checkout team at Total Wine & More, I check the tables behind checkout and billing first, because that is where a silent mismatch hurts most.
CDC vs dual writes: which is safer?
Dual writes mean your application writes to both the old and new database in the same code path. It looks simple, and it is tempting when you control the code.
The problem is that two databases cannot commit atomically without extra machinery. One write can succeed while the other fails, and now they disagree with no record of why. Concurrent requests can also apply changes to the two databases in different orders.
CDC avoids this by keeping one source of truth. The application writes once, and changes flow outward from the committed log. If the pipeline fails, it can resume from its last position and catch up, as long as the source still retains the log it needs.
| CDC | Dual writes | |
|---|---|---|
| Source of truth | One database | Two, which can drift |
| Changes to legacy code | Usually none | Required in every write path |
| Failure handling | Resume from log position | Custom reconciliation |
| Setup effort | Higher up front | Lower up front, higher later |
Dual writes can work when the system is small, you can see every write path and a brief mismatch is tolerable. Even then, I would pair them with a reconciliation job. For anything touching payments or user records, I would pick CDC first.
When is change data capture overkill?
CDC adds moving parts: a connector, a stream, consumers and monitoring. Someone has to run them. Skip it when:
- The database is small and a short maintenance window is acceptable.
- The data changes rarely, so a final export right before cutover catches everything.
- You only need a one-time copy, not continuous sync.
- The team has no one who can operate and monitor the pipeline after launch.
A simple rule of thumb from my own experience (an estimate, not a benchmark): if you can explain every write path on a whiteboard and downtime is cheap, try the simple option first. If users write constantly and downtime is not an option, CDC earns its cost.
Want a second opinion on your migration?
If you are planning a migration and are unsure whether CDC fits your case, email me@filipeeduardo.dev with a short description of your stack and write volume. I’m glad to compare notes and tell you what I would check first.
Frequently asked questions
Is CDC the same as database replication?
They are related. Native replication usually copies changes between databases of the same engine. CDC is more general: it exposes changes as events you can send to a different engine, a search index or a warehouse.
Does CDC slow down the source database?
Log-based CDC adds relatively little load because it reads a log the database already writes. Trigger-based CDC adds work to every write. Test under realistic traffic either way, and watch the log retention settings so unread logs do not fill the disk.
Can CDC handle deletes?
Log-based CDC can, since deletes appear in the log. Polling on a timestamp cannot see a deleted row, unless you use soft deletes.
How long should the two databases stay in sync?
Long enough to verify the new system and keep a rollback path, then no longer. Running both indefinitely creates cost and confusion, so set a date to retire the old one.
Do I still need a rollout plan with CDC?
Yes. CDC keeps the data aligned, but you still need a way to route traffic gradually and switch back, such as feature flags or a routing layer in front of both systems. My own rule is to put that switch in front of everything before writing any new code.

