如何实现嵌套数组过滤计数 按餐厅ID分组统计评论数据
MongoDB 嵌套评论数组分组统计实现
直接用MongoDB原生聚合管道就能实现需求,不需要在应用层做额外计算,具体实现如下:
现有数据结构
业务集合单条文档格式:
{ _id: ObjectId(''), restaurantId: ObjectId(''), orderId: ObjectId(''), reviews: [ { type: "food", tags: ["good food", "nice food"] }, { type: "pricing", tags: ["best price", "good price"] } ] }
聚合实现代码
用$facet分两个独立分支分别统计type计数、type+标签组合计数,避免数组拆分导致的计数重复问题:
db.restaurantReview.aggregate([ { $facet: { // 分支1:按餐厅分组,统计不同评论type的总数量 typeStats: [ { $unwind: "$reviews" }, { $group: { _id: { restaurantId: "$restaurantId", type: "$reviews.type" }, typeCount: { $sum: 1 } } }, { $group: { _id: "$_id.restaurantId", types: { $push: { k: "$_id.type", v: "$typeCount" } } } }, { $project: { types: { $arrayToObject: "$types" } } } ], // 分支2:按餐厅分组,统计type+标签组合的总数量 tagStats: [ { $unwind: "$reviews" }, { $unwind: "$reviews.tags" }, { $group: { _id: { restaurantId: "$restaurantId", typeTag: { $concat: ["$reviews.type", "-", "$reviews.tags"] } }, tagCount: { $sum: 1 } } }, { $group: { _id: "$_id.restaurantId", tags: { $push: { k: "$_id.typeTag", v: "$tagCount" } } } }, { $project: { tags: { $arrayToObject: "$tags" } } } ] } }, // 关联合并两个分支的统计结果 { $project: { merged: { $map: { input: "$typeStats", as: "typeItem", in: { $mergeObjects: [ "$$typeItem", { $first: { $filter: { input: "$tagStats", cond: { $eq: ["$$this._id", "$$typeItem._id"] } } } } ] } } } } }, { $unwind: "$merged" }, { $replaceRoot: { newRoot: "$merged" } }, // 处理没有标签统计的餐厅,返回空对象避免字段缺失 { $project: { _id: 1, types: 1, tags: { $ifNull: ["$tags", {}] } } } ])
返回结果格式
执行后输出结构完全匹配需求:
{ _id: ObjectId("对应餐厅的restaurantId值"), types: { food: 4, pricing: 2, ambience: 1 }, tags: { "food-good food": 3, "food-nice food": 2, "pricing-best price": 2, "ambience-superb": 1 } }
注意事项
- 两个统计分支独立做数组拆分,从根源上避免了同一条评论多标签导致的type计数虚高问题
- 如果需要调整标签的格式(比如把空格转驼峰、去掉特殊字符),可以在分支2的
$concat阶段前加$replaceOne/$replaceAll等字符串操作符处理 - 数据量较大时,提前给
restaurantId字段建普通索引,聚合速度会有明显提升
内容的提问来源于stack exchange,提问作者Abu Ra1han
相关产品推荐
相关产品推荐

