MongoDB聚合查询问题:排除指定title值执行分组聚合
MongoDB聚合查询:排除指定title并分组统计
你这段代码的问题出在$group阶段里的title :{$ne: "class notes"}写法错误——$group阶段的字段只能用聚合操作符(比如$sum、$first),不能直接用查询操作符做过滤,所以这行代码根本起不到排除"class notes"的作用,反而会给每个分组新增一个布尔值的title字段,完全不符合需求。
下面是两种正确的实现方式:
方案一:先过滤再分组(推荐,效率更高)
先通过$match排除所有title为"class notes"的记录,再进行分组统计,这样能减少后续分组的数据量,提升性能:
let coAuthorCutThreshold = 100; const cutTitles = []; network.aggregate([ // 先排除title为"class notes"的记录 { $match: { title: { $ne: "class notes" } } }, // 按title分组统计数量 { $group: { _id: '$title', count: { $sum: 1 } } }, // 筛选数量超过阈值的分组 { $match: { count: { $gt: coAuthorCutThreshold } } }, // 按数量降序排序 { $sort: { count: -1 } } ]).forEach(function(obj) { cutTitles.push(obj._id); });
方案二:分组后再过滤
如果需要先统计所有title的数量(包括"class notes"),再排除该分组,可以在分组后加一个$match过滤_id不等于"class notes":
let coAuthorCutThreshold = 100; const cutTitles = []; network.aggregate([ { $group: { _id: '$title', count: { $sum: 1 } } }, // 先排除"class notes"的分组,再筛选数量阈值 { $match: { _id: { $ne: "class notes" }, count: { $gt: coAuthorCutThreshold } } }, { $sort: { count: -1 } } ]).forEach(function(obj) { cutTitles.push(obj._id); });
两种方案都能实现你的需求,优先选方案一,因为提前过滤能减少聚合处理的数据量。
内容的提问来源于stack exchange,提问作者user110920
相关产品推荐
相关产品推荐

