You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助MongoDB从集合A复制唯一文档至集合B的查询语句

MongoDB 去重复制文档解决方案

步骤1:创建集合B的复合唯一索引

先在集合B上创建指定的复合唯一索引,从根源保证文档唯一性:

db.collectionB.createIndex(
  {
    "aC": 1,
    "aI": 1,
    "aN": 1,
    "aT": 1,
    "transaction.seqNo": 1,
    "transaction.tC": 1,
    "pD": 1
  },
  { unique: true }
);

步骤2:从集合A筛选唯一文档并复制到集合B

针对360万+条数据的规模,推荐用聚合管道结合$merge操作,在数据库层面高效完成去重与写入,避免客户端内存过载:

db.collectionA.aggregate([
  // 按唯一索引字段分组,保留每组首个文档
  {
    $group: {
      _id: {
        aC: "$aC",
        aI: "$aI",
        aN: "$aN",
        aT: "$aT",
        transactionSeqNo: "$transaction.seqNo",
        transactionTC: "$transaction.tC",
        pD: "$pD"
      },
      doc: { $first: "$$ROOT" }
    }
  },
  // 恢复原文档结构
  { $replaceRoot: { newRoot: "$doc" } },
  // 合并到集合B,重复文档直接跳过
  {
    $merge: {
      into: "collectionB",
      on: [
        "aC", "aI", "aN", "aT",
        "transaction.seqNo", "transaction.tC", "pD"
      ],
      whenMatched: "skip",
      whenNotMatched: "insert"
    }
  }
], { allowDiskUse: true }); // 大数据量下启用磁盘临时存储,避免内存不足

补充说明

  • $first可替换为$last,仅影响选取每组的首/尾文档,去重效果一致;
  • 若需覆盖已存在的重复文档,可将whenMatched改为replace;
  • allowDiskUse: true参数针对超大规模数据集,防止聚合因内存不足中断。

内容的提问来源于stack exchange,提问作者Shilpa1904

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 17:55:22