MongoDB近90天数据去重性能瓶颈问题求助
针对MongoDB近90天文档去重的性能优化方案
1. 优化索引,缩小数据扫描范围
- 为
created_at和fingerprint创建复合索引,让日期筛选和分组操作都能命中索引,避免全集合扫描:// C#驱动创建复合索引 var indexKeys = Builders<YourDocument>.IndexKeys.Ascending(d => d.created_at).Ascending(d => d.fingerprint); var indexModel = new CreateIndexModel<YourDocument>(indexKeys); await collection.Indexes.CreateOneAsync(indexModel); - 替换
$search为$match做日期范围筛选:$search是全文搜索能力,不适合日期范围过滤,$match可直接利用上述复合索引快速定位近90天数据,大幅减少后续处理量:var ninetyDaysAgo = DateTime.UtcNow.AddDays(-90); var matchStage = Builders<YourDocument>.Filter.Gt(d => d.created_at, ninetyDaysAgo);
2. 精简聚合管道,降低分组内存消耗
- 分组时只保留必需字段:不要
$push完整文档,仅保留_id即可——去重只需要知道待删除的文档ID,无需完整数据:
这种方式能大幅降低分组阶段的内存占用,同时满足你获取重复条目ID列表的需求。var aggregation = collection.Aggregate() .Match(matchStage) .Project(d => new { d._id, d.fingerprint }) // 仅投影必要字段 .Group( key => key.fingerprint, g => new { Fingerprint = g.Key, DuplicateIds = g.Select(d => d._id).ToList(), Count = g.Count() }) .Match(g => g.Count > 1);
3. 分批处理,避免一次性加载过量数据
- 将90天时间范围拆分为多个小批次(比如每7天一批),逐个批次处理,避免MongoDB因内存不足触发磁盘临时存储:
var startDate = DateTime.UtcNow.AddDays(-90); var endDate = DateTime.UtcNow; var batchDays = 7; while (startDate < endDate) { var batchEnd = startDate.AddDays(batchDays); if (batchEnd > endDate) batchEnd = endDate; var batchMatch = Builders<YourDocument>.Filter.And( Builders<YourDocument>.Filter.Gte(d => d.created_at, startDate), Builders<YourDocument>.Filter.Lt(d => d.created_at, batchEnd) ); // 执行当前批次的聚合逻辑 var batchDuplicates = await collection.Aggregate() .Match(batchMatch) .Project(d => new { d._id, d.fingerprint }) .Group( key => key.fingerprint, g => new { Fingerprint = g.Key, DuplicateIds = g.Select(d => d._id).ToList(), Count = g.Count() }) .Match(g => g.Count > 1) .ToListAsync(); // 处理当前批次的重复数据删除 await ProcessDuplicates(batchDuplicates, collection); startDate = batchEnd; }
4. 批量删除,提升去重执行效率
- 拿到重复ID列表后,使用
BulkWrite批量删除,比单条删除效率提升数倍:private async Task ProcessDuplicates(List<DuplicateGroup> duplicates, IMongoCollection<YourDocument> collection) { var bulkOperations = new List<WriteModel<YourDocument>>(); foreach (var group in duplicates) { // 保留第一个文档,删除其余重复项 var idsToDelete = group.DuplicateIds.Skip(1).ToList(); if (idsToDelete.Count == 0) continue; var deleteFilter = Builders<YourDocument>.Filter.In(d => d._id, idsToDelete); bulkOperations.Add(new DeleteManyModel<YourDocument>(deleteFilter)); } if (bulkOperations.Count > 0) { await collection.BulkWriteAsync(bulkOperations); } } // 分组结果对应的实体类 public class DuplicateGroup { public string Fingerprint { get; set; } public List<ObjectId> DuplicateIds { get; set; } public int Count { get; set; } }
5. 可选:用临时集合中转聚合结果
- 如果聚合结果数据量仍较大,可先将分组后的重复数据写入临时集合,再从临时集合读取执行删除,避免聚合管道长时间占用资源:
// 将聚合结果写入临时集合 await collection.Aggregate() .Match(matchStage) .Project(d => new { d._id, d.fingerprint }) .Group( key => key.fingerprint, g => new { Fingerprint = g.Key, DuplicateIds = g.Select(d => d._id).ToList(), Count = g.Count() }) .Match(g => g.Count > 1) .OutAsync("temp_duplicates"); // 从临时集合读取并处理删除 var tempCollection = database.GetCollection<DuplicateGroup>("temp_duplicates"); var allDuplicates = await tempCollection.Find(_ => true).ToListAsync(); await ProcessDuplicates(allDuplicates, collection); // 清理临时集合 await tempCollection.DropAsync();
内容的提问来源于stack exchange,提问作者samivagyok
相关产品推荐
相关产品推荐

