Change Data Capture: Building Reliable Real-Time Data Pipelines Without Breaking Production
A practical guide to CDC architecture, Debezium, AWS DMS, transactional outbox, idempotency, schema evolution, monitoring, replay, and near-real-time data pipelines.
By BoundLayer Engineering Team
BoundLayer is a senior engineering partner for SaaS, fintech, AI automation, cloud infrastructure, legacy modernization, Web3, IoT, GPU computing, and data systems.
Many companies reach the same architectural problem: important business data is trapped inside operational databases, while analytics, automation, search, AI systems, and customer-facing products need a fresh copy of it somewhere else.
The first solution is usually a scheduled export. It is simple, understandable, and often correct. As the business grows, however, hourly or nightly jobs start creating stale dashboards, delayed alerts, duplicated processing, expensive full-table scans, and increasingly fragile integrations.
Change Data Capture, usually shortened to CDC, offers a different model. Instead of repeatedly copying every row, a CDC pipeline reads committed database changes and turns inserts, updates, and deletes into a continuous stream.
That stream can keep a warehouse current, update a search index, synchronize services, trigger business workflows, or feed an operational AI system. But CDC is not automatic correctness. A production design must handle ordering, duplicates, schema changes, recovery, backfills, privacy, and the operational impact on the source database.
At BoundLayer, we build data and integration systems around explicit business latency and reliability requirements. This guide explains where CDC creates value, how to design it safely, and when a simpler batch process remains the better choice.
What Change Data Capture Actually Captures
A transactional database records changes so it can recover after failure and replicate state. Depending on the database, this history may live in a write-ahead log, binary log, redo log, or another engine-specific mechanism.
A log-based CDC connector reads that ordered history and emits an event for each committed change.
A typical event contains:
- the source database, schema, table, and primary key;
- the operation type: create, update, or delete;
- previous and current row values when available;
- the source transaction position;
- transaction and commit timestamps;
- connector metadata and schema version.
The important distinction is that CDC observes the database's committed state transitions. It does not require application code to remember to publish a second message after writing a row.
Application
|
Transactional database
|
Transaction log
|
CDC connector
|
Event stream or durable storage
|
Warehouse | Search | Cache | Services | Automation | AI
Tools such as Debezium, cloud database replication services, and database-native logical replication implement variations of this architecture.
CDC Is Usually Near Real Time, Not Instantaneous
The phrase "real time" is frequently used without a measurable requirement.
A CDC event must be written by the source, read by a connector, transmitted through infrastructure, processed by one or more consumers, and committed to a target. Each stage adds delay and can accumulate a backlog.
AWS explicitly notes in its Database Migration Service CDC documentation that replication latency depends on source workload, network conditions, replication resources, target capacity, and data characteristics. There is no universal promise of instantaneous replication.
Define freshness as a service objective:
- 99% of inventory changes reach search within 15 seconds;
- finance dashboards are no more than five minutes behind the ledger;
- a customer status change reaches downstream workflows within one minute;
- an offline feature table is updated before the next scoring window.
This turns "we need real time" into an architecture and capacity question that can be tested.
Where CDC Creates Business Value
Operational analytics
Executives and operations teams can see orders, payments, inventory, or service activity without running analytical queries against the production application database.
Search and materialized views
Product, document, or customer changes can update a search engine or read-optimized store. The application keeps transactional ownership while the target is designed for fast queries.
Event-driven automation
A committed business change can trigger fulfillment, notifications, risk checks, account provisioning, or internal workflows. This is useful when a legacy application cannot easily publish domain events itself.
Cache invalidation
Consumers can invalidate or refresh cached records after database changes instead of relying only on time-to-live expiration.
Zero- or low-downtime migration
A team can perform an initial load, continuously replicate new changes, validate the target, and switch traffic after lag approaches an acceptable threshold. AWS provides a reference architecture for using DMS with Kinesis or Amazon MSK for continuous change ingestion.
AI and retrieval systems
AI assistants and agents need current, governed business context. CDC can update indexes, feature stores, entity views, and workflow state without repeatedly scanning source systems. It complements the retrieval and live-tool patterns described in our guide to context engineering for AI agents.
CDC, Polling, Batch, and Application Events
CDC is one option among several integration patterns.
| Pattern | Strengths | Limitations |
|---|---|---|
| Scheduled batch | Simple, cheap, easy to replay | Stale data, repeated scans, coarse recovery |
| Timestamp polling | Easy for small systems | Missed changes, clock issues, difficult deletes |
| Micro-batch | Good balance of freshness and simplicity | Not event-level latency |
| Log-based CDC | Efficient incremental capture, includes deletes | Operational complexity and source-specific behavior |
| Application events | Rich business meaning | Requires application changes and reliable publication |
| Dual writes | Appears simple initially | Can leave database and message system inconsistent |
CDC events describe data changes. They do not automatically describe business intent.
An update to orders.status may mean that an order was paid, cancelled, corrected by support, or changed during a migration. A domain event such as OrderPaymentConfirmed carries richer semantics.
Use CDC when downstream systems need faithful database changes or when modifying the source application is impractical. Prefer explicit application events when consumers need stable business meaning. Many mature systems use both.
The Transactional Outbox Pattern
The transactional outbox combines application intent with reliable CDC delivery.
Instead of writing business state to the database and separately publishing a message, the application writes both the state change and an outbox record in one database transaction. A CDC connector then publishes the outbox record.
BEGIN TRANSACTION
update order state
insert OrderPaymentConfirmed into outbox
COMMIT
|
CDC connector
|
event broker
If the transaction commits, both records exist. If it rolls back, neither exists. This removes the inconsistency window created by a non-transactional dual write.
Debezium documents its outbox event router as a way to reliably exchange data between services while keeping internal database state and published events consistent.
An outbox record normally includes:
- a globally unique event ID;
- aggregate type and aggregate ID;
- event type and schema version;
- creation time;
- serialized payload;
- optional tenant, trace, or correlation metadata.
Consumers must still be idempotent because message delivery can be repeated.
A Production CDC Reference Architecture
A production pipeline needs distinct capture, transport, processing, storage, and control responsibilities.
PostgreSQL / MySQL / SQL Server / Oracle
|
Log-based CDC connector
|
Durable event transport
| | |
Raw archive Stream jobs Dead-letter path
| |
Replay source Validated data products
| | |
Warehouse Search Business workflows
Control plane: schemas, checkpoints, deployment, access, monitoring
Source database
Configure the database log, retention, replication identity, and least-privilege connector account. Confirm how large transactions, table rewrites, and DDL operations behave.
Connector
The connector reads source positions, serializes changes, persists checkpoints, and resumes after interruption. It needs clear failure and upgrade procedures.
Durable transport
Kafka, Amazon MSK, Kinesis, or another durable stream decouples capture from consumers. Partitioning, retention, throughput, and access control determine whether the pipeline can recover safely.
Raw archive
An immutable object-store copy gives the team a longer replay window, audit evidence, and a source for rebuilding targets after consumer defects.
Transformation and validation
Consumers normalize source events into stable data products or domain events. This is where schema enforcement, filtering, enrichment, deduplication, and data-quality checks belong.
Targets
Warehouses, lakehouses, search indexes, caches, feature stores, and application services each need target-specific idempotency and reconciliation logic.
Delivery Semantics: Design for Duplicates
Most practical CDC pipelines provide at-least-once delivery across the complete path. A connector or consumer can process an event, fail before saving its checkpoint, and process it again after restart.
That is not an exceptional corner case. It is a normal recovery scenario.
Consumers should use one or more of these strategies:
- upsert by stable primary key;
- store and reject previously processed event IDs;
- compare monotonically increasing source positions or versions;
- write the target result and consumer checkpoint atomically;
- make side effects conditional on a unique idempotency key.
Do not describe a system as exactly once merely because one component supports transactions. End-to-end guarantees include the source, connector, broker, processor, target, and external side effects.
For email, payment, or provisioning actions, a duplicate has business consequences. Persist an idempotency record before or atomically with the action whenever the external system permits it.
Ordering and Partitioning
A database transaction log has an order, but a distributed stream cannot preserve one global order at unlimited scale.
Events are usually partitioned by a stable key such as customer, account, order, or aggregate ID. Events for that key remain ordered within one partition while different keys are processed concurrently.
Choose the key according to the invariant that must be protected.
If every event for one order must be applied sequentially, partition by order ID. If updates across an entire account must remain ordered, account ID may be correct but can create hot partitions for large tenants.
Also decide how to handle:
- transactions that modify multiple tables;
- events arriving from several source databases;
- late events after retry;
- target writes that complete out of order;
- repartitioning during topology changes.
Global ordering is expensive and often unnecessary. Define the smallest business scope that requires order.
Initial Loads, Backfills, and Cutover
A new CDC pipeline normally needs both historical state and ongoing changes.
A safe bootstrap process is:
- Record or establish a source log position.
- Take a consistent snapshot or full load.
- Start or resume change capture from the recorded position.
- Apply buffered changes to the target.
- Reconcile source and target counts, checksums, and business totals.
- Expose the target only after freshness and correctness meet the cutover criteria.
Backfills are not the same as live events. A backfill can overwhelm consumers, reorder updates, or overwrite a newer target value with an older snapshot.
Use separate capacity and explicit version comparison. Tag backfill records, throttle them, and make the merge rule deterministic.
For migration work, this technique fits a broader staged approach that avoids risky one-time cutovers. See our guide to legacy-to-cloud modernization for the surrounding discovery, stabilization, migration, and rollback process.
Schema Evolution Is an Operational Contract
Source teams will rename columns, change types, add constraints, and deploy new application versions. If downstream consumers discover those changes only after failure, the organization does not have a reliable data platform.
Establish a schema policy:
- additive fields are backward compatible by default;
- removals and renames require a deprecation window;
- incompatible type changes create a new field or event version;
- schemas are stored and validated centrally;
- consumers declare compatible versions;
- CI checks proposed database migrations against downstream contracts;
- ownership and escalation are visible.
Raw table CDC tightly couples consumers to storage details. A normalization layer can translate source-specific envelopes into stable contracts, but that layer itself becomes critical infrastructure.
Treat schema evolution as product lifecycle management, not only serialization configuration.
Deletes, Privacy, and Data Retention
Deletes are one reason timestamp polling fails. A row that no longer exists cannot appear in a query of current rows.
CDC can emit a delete event or tombstone, but each target must define what deletion means:
- remove the document from search;
- mark an analytical row as deleted;
- invalidate a cache entry;
- retain an audit record under a legal basis;
- propagate an erasure request to derived datasets.
Replicating data creates additional copies, and every copy expands the security and privacy boundary.
Apply data minimization before broad distribution. Filter unused columns, tokenize identifiers where possible, encrypt transport and storage, restrict topic access, audit consumption, and define retention for streams and archives. Sensitive payloads should not be copied into logs or dead-letter queues without equivalent controls.
Monitoring the Pipeline End to End
Connector health alone does not prove that downstream data is correct or fresh.
Monitor four layers.
Capture health
- connector status and restarts;
- source log position and retained-log pressure;
- replication slot or equivalent state;
- transaction size and capture errors.
Transport health
- publish failures;
- partition throughput and skew;
- retention headroom;
- broker storage and replication health.
Consumer health
- consumer lag and oldest unprocessed event age;
- processing latency, retry rate, and dead-letter volume;
- schema validation failures;
- duplicate and out-of-order handling.
Data health
- source-to-target freshness;
- row counts and sampled checksums;
- business totals, such as order value by day;
- missing keys and referential anomalies;
- delete propagation and reconciliation failures.
Alert on business freshness, not only infrastructure availability. A healthy connector feeding a failed transformation still produces stale customer data.
Recovery and Replay
Every CDC design should answer these questions before launch:
- What is the last durable source position?
- How long are source logs retained?
- How long are events retained in the broker and archive?
- Can a single consumer replay without affecting others?
- How is a corrected transformation applied to historical events?
- What happens if the target is unavailable for hours?
- How are poison events isolated without silently losing data?
Use dead-letter handling as a quarantine, not a disposal mechanism. Store enough metadata to diagnose and replay a rejected event, assign ownership, and alert before the queue becomes an ignored data cemetery.
Periodically rehearse recovery. A replay process that has never been tested is only a document.
When CDC Is the Wrong Choice
Streaming infrastructure adds connectors, brokers, schemas, checkpoints, operational alerts, and distributed failure modes.
Do not adopt it only because "real time" sounds modern.
Batch or micro-batch is usually better when:
- the business consumes data daily or hourly;
- source volume is small;
- a five- or fifteen-minute delay has no measurable impact;
- the target can be rebuilt cheaply;
- the team cannot operate streaming infrastructure;
- a managed incremental connector already meets the requirement;
- consumers need only final aggregates, not every state transition.
A useful decision test is the cost of stale data. If reducing freshness from one hour to ten seconds does not change a decision, customer experience, risk exposure, or operating process, the extra system may not be justified.
Micro-batching often provides the best middle ground: bounded delay, simpler recovery, efficient bulk writes, and lower operational overhead.
AWS Implementation Options
AWS offers several building blocks, and the correct combination depends on whether the primary goal is migration, analytics, or application events.
AWS Database Migration Service
DMS supports full loads and ongoing replication from multiple relational and non-relational sources. It is useful for migrations and managed change delivery, but teams must benchmark actual lag, supported data types, DDL behavior, and recovery requirements.
Amazon MSK or Amazon Kinesis
Use a durable stream when multiple consumers need independent processing, retention, scaling, and replay. Kafka ecosystems provide broad connector and stream-processing capabilities; Kinesis reduces some operational ownership within AWS.
Amazon S3
An S3 raw archive provides durable, cost-efficient history for audit and replay. Partitioning and file-size management matter for downstream query cost.
Lambda, containers, or stream processors
Simple transformations may fit Lambda. Stateful joins, windows, high throughput, or complex replay can require Flink, Kafka Streams, Spark, or custom containerized consumers.
Redshift, OpenSearch, and purpose-built stores
Targets should match query and product requirements. Do not use the stream itself as the only queryable system of record.
An AWS infrastructure audit should include the networking, encryption, IAM, observability, multi-AZ strategy, retention cost, and disaster-recovery behavior around these components.
A Practical Delivery Plan
Phase 1: Define the outcome
Identify consumers, freshness objectives, data ownership, volume, security classification, and the business cost of delay or data loss.
Phase 2: Inspect the source
Review database version, transaction log configuration, primary keys, large objects, transaction patterns, DDL process, load headroom, and maintenance ownership.
Phase 3: Build one bounded pipeline
Choose a small but valuable dataset. Implement capture, durable transport, idempotent consumption, schema validation, monitoring, reconciliation, and replay.
Phase 4: Load and validate history
Run a controlled snapshot and catch-up process. Compare technical counts and business invariants before cutover.
Phase 5: Production hardening
Load-test bursts and large transactions, simulate connector and target failures, verify log-retention headroom, test schema changes, and rehearse recovery.
Phase 6: Expand through reusable standards
Create templates for connectors, topics, schemas, metrics, alerts, permissions, retention, and consumer libraries. Scale the operating model, not only the number of pipelines.
This is a good fit for forward deployed engineering: the work crosses application teams, database ownership, cloud infrastructure, analytics, and business operations, so it benefits from senior engineers working directly with stakeholders through production adoption.
How BoundLayer Can Help
BoundLayer designs and implements production data pipelines, integrations, cloud systems, and workflow automation.
We can help you:
- assess whether CDC, micro-batch, or scheduled batch fits the business requirement;
- design AWS DMS, Debezium, Kafka, MSK, Kinesis, and S3 architectures;
- implement transactional outbox and reliable domain-event publishing;
- migrate databases with controlled replication and cutover;
- build idempotent consumers for warehouses, search, caches, and business systems;
- introduce schema governance, data-quality checks, reconciliation, and replay;
- secure sensitive data across streams, archives, and downstream stores;
- diagnose lag, duplication, data loss, hot partitions, and fragile pipelines;
- connect fresh operational data to automation and AI systems.
The goal is not to add streaming technology everywhere. It is to deliver the right data at the required time, with failure behavior the business can understand and operate.
Final Takeaway
Change Data Capture is a powerful bridge between transactional systems and modern data products. It can reduce source load, improve freshness, support migrations, and unlock event-driven automation.
Its value comes from disciplined system design: explicit freshness objectives, durable checkpoints, idempotent consumers, controlled schemas, reconciliation, security, and tested recovery.
Start with the business decision that fresher data will improve. Then build the smallest reliable pipeline that satisfies it. When near-real-time data is genuinely valuable, CDC becomes infrastructure for better operations rather than another source of distributed complexity.
Related engineering articles
AWS Infrastructure Audit and Optimization: How We Build, Fix, and Scale Cloud Platforms
How BoundLayer audits, creates, secures, optimizes, and operates AWS infrastructure for SaaS, fintech, AI, data, and backend-heavy products.
Legacy-to-Cloud Modernization: Building a More Reliable, Scalable, and Cost-Efficient System
How BoundLayer modernizes legacy applications for the cloud: discovery, stabilization, migration strategies, cost control, reliability, security, and measurable outcomes—without risky full rewrites.
GPU Inference Optimization: How to Reduce AI Model Serving Cost Without Sacrificing Reliability
A practical guide to reducing GPU inference cost with model right-sizing, quantization, batching, caching, autoscaling, observability, and reliable production architecture.
Need reliable, fresh data across your systems?
We design and build CDC, streaming, migration, and integration pipelines with measurable freshness, data quality, security, and recovery controls.