Confluent Connect-JDBC与精确一次交付:自增主键场景下是否防丢行重复?
Great question—this is a super common scenario when setting up JDBC source connectors, so let’s break down how it works and what you need to configure to get reliable delivery.
Avoiding Data Loss: Yes, Under Normal Circumstances
When using an auto-increment primary key as your incremental field (via mode=incrementing), the connector operates by tracking the highest value of this key it’s successfully processed. On each poll, it only fetches rows where the primary key is greater than this tracked value.
- Since auto-increment keys are monotonically increasing (most databases like MySQL or PostgreSQL don’t reuse auto-increment values even if rows are deleted), every new inserted row will have a higher key than the last. This means the connector won’t miss any new records as long as:
- You don’t manually modify the auto-increment key to a lower value (a practice that’s highly discouraged in production).
- The database’s auto-increment mechanism isn’t broken or misconfigured.
The only edge case for data loss would be if you truncate the table and reset the auto-increment counter, but that’s a deliberate destructive operation—not a normal runtime scenario.
Avoiding Duplicate Rows: Depends on Exactly Once Configuration
By default, the incrementing mode operates in an at-least-once delivery model. Here’s why: if the connector crashes after fetching rows but before committing its offset to Kafka, it will re-fetch those same rows when it restarts, leading to duplicates in your Kafka topic.
To eliminate duplicates and achieve exactly-once delivery, you need to enable and configure the right settings:
- Enable exactly-once support in the connector: Set
exactly.once.support=true(note: configuration names may vary slightly by Confluent Platform version—check your docs for the exact key). This requires your Kafka cluster to support transactions (broker version ≥ 0.11.0, withtransaction.state.log.replication.factorset to at least 3). - Ensure proper database isolation: Use a transaction isolation level like
READ COMMITTEDor higher to prevent the connector from reading uncommitted rows that might later be rolled back, which could lead to re-processing. - Pair with transaction-aware consumers: Even if the connector produces exactly-once, your downstream consumers need to support transactional reads (using
isolation.level=read_committed) to avoid processing duplicates from Kafka’s transaction logs.
Critical Caveats
- Don’t update auto-increment keys: Auto-increment primary keys should never be updated (it’s against best practices anyway). The
incrementingmode only tracks new inserts—updates to the key won’t be detected, leading to data inconsistency (though not loss or duplicates). - Prevent key overflow: Choose a sufficiently large data type for your auto-increment key (e.g.,
BIGINTinstead ofINTin MySQL) to avoid hitting the maximum value and causing key collisions, which would break the incremental tracking logic.
Final Verdict
With proper configuration (enabling exactly-once support, maintaining a valid auto-increment primary key, and using transaction-aware components end-to-end), Kafka Connect-JDBC can both avoid data loss and prevent duplicate rows when using an auto-increment primary key as the incremental field.
内容的提问来源于stack exchange,提问作者glarus089

