MongoDB高效统计指定日期文档resources去重值数量的方案
高效统计MongoDB中resources字段的去重值数量
问题背景
现有MongoDB文档结构如下:
{ "date": "2023-11-09", "type": "my_type", "resources": "1111,5555,2222,3333,1111" }
需求:获取指定日期下的所有文档(最多50万条),统计resources字段中逗号分隔值的去重总数。当前使用的聚合查询耗时超15秒,需要基于Reactive Spring Boot实现更高效的方案。
当前查询的问题
当前聚合查询存在逻辑错误且性能低下:
[ { $match: { date: "2024-01-01" } }, { $project: { resources: "$resources" } }, { $unwind: { path: "$resources" } }, { $group: { _id: null, dv: { $addToSet: "$resources" } } }, { $project: { total: { $size: "$dv" } } } ]
注:
$unwind仅对数组字段生效,当前resources是字符串,此步骤要么报错,要么将整个字符串作为单个元素处理,统计结果完全不符合预期,同时无意义的操作也会增加查询耗时。
优化方案
1. 数据结构重构(核心优化)
将resources从逗号分隔的字符串改为数组类型,存储为:
{ "date": "2023-11-09", "type": "my_type", "resources": ["1111","5555","2222","3333","1111"] }
这种结构可以直接利用MongoDB的数组操作符,避免字符串拆分的额外开销,大幅提升性能。如果无法直接修改历史数据,可在查询时临时拆分,但长远来看建议批量更新数据结构。
2. 添加索引加速查询
为date字段创建单字段索引,减少$match阶段的数据扫描范围:
db.your_collection.createIndex({ date: 1 })
如果查询时同时用到type字段,可创建复合索引:db.your_collection.createIndex({ date: 1, type: 1 })
3. 优化聚合查询逻辑
针对数组类型的resources(推荐)
直接通过$unwind拆分数组,再用$addToSet去重统计:
[ { $match: { date: "2024-01-01" } }, { $unwind: "$resources" }, { $group: { _id: null, uniqueResources: { $addToSet: "$resources" } } }, { $project: { total: { $size: "$uniqueResources" } } } ]
针对无法修改的字符串类型resources
先使用$split将字符串拆分为数组,再进行后续处理:
[ { $match: { date: "2024-01-01" } }, { $project: { resources: { $split: ["$resources", ","] } } }, { $unwind: "$resources" }, { $group: { _id: null, uniqueResources: { $addToSet: "$resources" } } }, { $project: { total: { $size: "$uniqueResources" } } } ]
注:此方案性能仍不如数组类型,因为每个文档都要执行字符串拆分操作。
4. Reactive Spring Boot中的实现
使用Spring Data MongoDB Reactive的AggregationAPI构建查询,示例代码:
import org.springframework.data.mongodb.core.ReactiveMongoTemplate; import org.springframework.data.mongodb.core.aggregation.Aggregation; import reactor.core.publisher.Mono; // 自定义结果映射类 class UniqueCountResult { private Long total; public Long getTotal() { return total; } public void setTotal(Long total) { this.total = total; } } public Mono<Long> countUniqueResources(String targetDate) { Aggregation aggregation = Aggregation.newAggregation( Aggregation.match(Criteria.where("date").is(targetDate)), Aggregation.unwind("resources"), Aggregation.group().addToSet("resources").as("uniqueResources"), Aggregation.project().andExclude("_id").and("uniqueResources").size().as("total") ); return reactiveMongoTemplate.aggregate(aggregation, "your_collection_name", UniqueCountResult.class) .map(UniqueCountResult::getTotal) .singleOrEmpty(); }
额外性能优化建议
- 如果数据量接近50万条,可考虑分批次统计后合并结果,避免单次聚合占用过多内存。
- 定期清理过期数据,减少
$match阶段扫描的数据量。
内容的提问来源于stack exchange,提问作者Moussa ELaQ
相关产品推荐
相关产品推荐

