You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mongo中如何通过聚合查询product_id与orders数组项均重复的文档

MongoDB 筛选product_id与orders数组均重复文档实现方案

实现思路

核心逻辑是为每个文档生成唯一标识product_id + orders数组内容的特征签名,按签名分组后筛选计数≥2的分组,即可得到所有符合重复条件的文档。

具体实现方案

场景1:要求orders数组元素顺序、内容完全一致才判定为重复

直接对原始product_id和orders数组生成哈希签名,聚合语句如下:

db.your_collection_name.aggregate([
  // 生成唯一特征签名
  {
    $set: {
      unique_sign: {
        $md5: {
          $concat: [
            "$product_id",
            { $toString: "$orders" }
          ]
        }
      }
    }
  },
  // 按签名分组统计重复项
  {
    $group: {
      _id: "$unique_sign",
      repeat_count: { $count: {} },
      repeat_doc_ids: { $push: "$_id" },
      product_id: { $first: "$product_id" },
      orders: { $first: "$orders" }
    }
  },
  // 过滤出存在重复的分组
  {
    $match: {
      repeat_count: { $gte: 2 }
    }
  },
  // 可选:需要返回完整重复文档时添加该阶段
  {
    $lookup: {
      from: "your_collection_name",
      localField: "repeat_doc_ids",
      foreignField: "_id",
      as: "full_repeat_docs"
    }
  }
])

场景2:orders数组不要求顺序,只要元素集合完全一致即判定为重复

先对orders数组按统一规则排序后再生成签名,将第一个$set阶段替换为如下逻辑即可:

{
  $set: {
    // 对orders数组按row升序、seat升序排序,消除顺序影响
    sorted_orders: {
      $sortArray: {
        input: "$orders",
        sortBy: { row: 1, seat: 1 }
      }
    },
    unique_sign: {
      $md5: {
        $concat: [
          "$product_id",
          { $toString: "$sorted_orders" }
        ]
      }
    }
  }
}

大规模集合性能优化建议

  • 提前为product_id字段建立索引,降低聚合阶段的扫描开销
  • 不需要获取完整重复文档时,删除最后的$lookup阶段,仅保留重复文档ID列表即可,性能提升明显
  • 若使用的MongoDB版本低于5.2无$sortArray算子,可通过$function自定义JS函数实现数组排序,或在数据写入阶段就固定orders数组的元素排列顺序

内容的提问来源于stack exchange,提问作者Malmoc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 05:15:02