基于复制的分布式数据库能否真正实现数据删除?
Great question—let’s unpack your core assumption and break down the reality of how deletion works in these systems:
此前我一直认为,在基于复制的分布式数据库中无法真正删除行数据,仅能在基于拷贝的数据库中正常完成该操作。在复制架构中,仅会将数据标记为“待删除”并在所有查询中过滤,并未实际从数据库中删除数据。
Why Tombstones (Marked-for-Deletion) Are Common in Replication Systems
You’re absolutely correct that many replication-based distributed databases rely on tombstones (entries marked as deleted) instead of immediate physical deletion. The main reasons are:
- Replication lag protection: If a node physically deletes a row before the delete command propagates to all replicas, a query hitting an out-of-sync replica would return stale data, breaking consistency.
- Eventual consistency guarantees: Most replication systems prioritize availability and partition tolerance over strict consistency. Tombstones ensure that even if a replica misses the delete event initially, it will process the tombstone later and filter the data out.
Exceptions: When Physical Deletion Does Happen
That said, your assumption isn’t universally true—some replication-based databases do support actual physical deletion under specific conditions:
- Synchronous replication with strict consistency: Systems that wait for all replicas to acknowledge the delete before committing (like PostgreSQL with synchronous streaming replication) can safely perform physical deletion without consistency risks.
- Background compaction: Databases like Cassandra or DynamoDB use background compaction processes to permanently remove tombstones after a configured retention period. So while deletion starts as a tombstone, the data is eventually physically erased from storage.
Race Conditions in Key-Value Replication
Your hunch about key-value systems introducing race conditions during replication is spot-on. A classic scenario looks like this:
- Node A receives a delete for key
Xand creates a tombstone. - Before the tombstone replicates to Node B, Node B gets an update for key
X. - Without proper conflict resolution (like vector clocks or timestamp ordering), the update can overwrite the tombstone, bringing the "deleted" data back—this is a critical race condition that replication systems need to handle explicitly.
How to Verify Your Hypothesis
To put this to the test, you could:
- Pick a specific replication-based database (e.g., PostgreSQL with streaming replication, Cassandra) and run deletion tests.
- Inspect the underlying storage files before deletion, right after deletion, and after compaction to see if row data is physically removed.
- Simulate replication lag (using tools to throttle network traffic between nodes) and test if queries return consistent results post-deletion.
内容的提问来源于stack exchange,提问作者Christopher

