RocksDB:CompactOnDeletionCollector触发压缩的时机、优先级及配置无效果问题
Let's break down your questions one by one, and figure out why your test isn't showing the expected differences.
1. When does compaction trigger after CompactOnDeletionCollector marks an SST file?
The CompactOnDeletionCollector doesn't trigger compaction directly—it adds a metadata flag to SST files when they meet your configured deletion thresholds. Here's when RocksDB will act on that flag:
- During regular background checks: RocksDB's compaction scheduler periodically scans for SST files needing cleanup. When it encounters a file marked by this collector, it will prioritize adding it to the compaction queue.
- After memtable flushes: When a memtable is flushed to disk as an SST, RocksDB will check the new file (and nearby files) for the deletion flag and consider them for compaction.
- On explicit compaction triggers: Operations like manual
CompactRangecalls, or when L0 file counts hit your trigger threshold, will also include flagged files in compaction candidates.
The key takeaway: the collector doesn't start compaction from scratch—it makes eligible files get picked up faster by RocksDB's existing compaction machinery.
2. What's the priority of this compaction vs regular compaction?
Files marked by CompactOnDeletionCollector get higher priority than regular compaction candidates. RocksDB's compaction picker logic prioritizes these files because they have a high ratio of deleted data (or a large absolute volume of deletions), so compacting them recovers wasted storage space more efficiently.
In practice, if there are both flagged SST files and files eligible for regular compaction (e.g., L0 file count hitting the trigger), the flagged files will be processed first by background compaction threads.
3. Why isn't your test showing any differences?
Your test setup has a few critical issues preventing the collector from kicking in:
a. DeleteRange creates a single tombstone, not 100k individual deletions
Your test uses DeleteRange to remove all 100k keys, but RocksDB stores range deletions as a single tombstone entry (not one per key). The CompactOnDeletionCollector counts individual deletion entries (or deletion bytes) to trigger the flag. Your configured threshold is min_delete_count=10000, but each SST file only has 1 deletion entry (the range tombstone)—way below the threshold. So no files get marked for priority compaction.
To hit the threshold, you could either:
- Perform 10k+ individual
Deletecalls instead of a singleDeleteRange - Adjust your collector thresholds (e.g., lower
min_delete_countto 1, or tweakdeletion_ratio_thresholdto a value the range tombstone's byte count meets)
b. You're checking stats immediately before background compaction runs
RocksDB's compaction runs in background threads. Your test calls DeleteRange and immediately fetches stats then exits—there's no time for compaction to actually execute. Even if the collector had marked files, you wouldn't see stats updates yet.
Fix this by:
- Adding a short sleep after deletion (e.g.,
std::this_thread::sleep_for(std::chrono::seconds(5));) to give background threads time to run - Forcing a manual compaction with
db->CompactRange(rocksdb::CompactRangeOptions(), nullptr, nullptr);before fetching stats (though this bypasses the collector's priority logic, it will show you the cleanup effect)
c. Your compaction triggers might mask the collector's effect
Your config sets level0_file_num_compaction_trigger=2 and write_buffer_size=1MB. Writing 100k keys will generate multiple L0 files, which trigger regular compaction to L1 before you even run the deletion. By the time you delete, most data is already in L1, and the range tombstone might be in a small L0 file that doesn't hit your collector thresholds.
内容的提问来源于stack exchange,提问作者jeffreyveon

