如何在MongoDB中对$project生成的自定义字段执行$sum分组统计
问题根因
你的代码错误点出在$group阶段的_id赋值:你填写的是固定字符串"convert",而非引用上游管道生成的convert字段,MongoDB中引用字段需要添加$前缀,即写成"$convert"。
修复后代码(基于你的原有写法调整)
db.Tweets.aggregate([ { $project: { _id: 0, convert: { $cond: { if: { $or: [ { $regexMatch: { input: "$tweet", regex: /Tesla/i } }, { $regexMatch: { input: "$tweet", regex: /rocket/i } }, { $regexMatch: { input: "$tweet", regex: /coffee/i } } ] }, then: "contains words", else: "does not contain words" } } } }, { $group: { _id: "$convert", Number_Of_Tweets_Sent: { $sum: 1 } } } ])
更高效的简化写法(省去$project阶段)
可以直接将分类逻辑放在$group阶段中计算,减少管道处理步骤,提升查询性能:
db.Tweets.aggregate([ { $group: { _id: { $cond: { if: { $or: [ { $regexMatch: { input: "$tweet", regex: /Tesla/i } }, { $regexMatch: { input: "$tweet", regex: /rocket/i } }, { $regexMatch: { input: "$tweet", regex: /coffee/i } } ] }, then: "contains words", else: "does not contain words" } }, Number_Of_Tweets_Sent: { $sum: 1 } } } ])
注:上述代码补充了coffee正则的大小写不敏感标识i,和另外两个正则的匹配规则保持一致,你可以根据实际需求调整。
内容的提问来源于stack exchange,提问作者Christopher
相关产品推荐
相关产品推荐

