You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在MongoDB中基于公共字段合并多集合并生成统一集合?

MongoDB 多集合公共字段合并并标注源集合

我的数据库里有col1、col2、col3三个集合,分别存储不同事件的用户信息,文档字段各不相同,但都有full_name、email_address这些公共字段。我想创建一个新集合outcol,只保留所有集合共有的字段数据,同时标注每条数据来自哪个源集合,以此整合所有公共数据。试过用aggregate函数,但没找到合适的多集合合并方法(大多资料只讲基于公共字段关联集合)。


示例集合插入代码

col1 插入代码

db.col1.insertMany([
  {
    "full_name": "Joel K",
    "email_address": "joelk@gmail.com",
    "mobile": "9343658663",
    "date": "February 5, 2022",
    "time": "11:59 AM",
    "company": "Apple Inc",
    "designation": "Manager"
  },
  {
    "full_name": "Sam K",
    "email_address": "samk@gmail.com",
    "mobile": "9343858663",
    "date": "February 7, 2022",
    "time": "09:59 AM",
    "company": "Google",
    "designation": "COO"
  }
])

col2 插入代码

db.col2.insertMany([
  {
    "full_name": "Karan H",
    "email_address": "karanh@gmail.com",
    "mobile": "93436545323",
    "date": "February 6, 2022",
    "message": "Welcome"
  },
  {
    "full_name": "John H",
    "email_address": "john@gmail.com",
    "mobile": "9343854536",
    "date": "February 9, 2022",
    "message": "Hello"
  }
])

col3 插入代码

db.col3.insertMany([
  {
    "full_name": "Tristan J",
    "email_address": "tristanj@gmail.com",
    "time":"03:00 PM"
  },
  {
    "full_name": "Richard S",
    "email_address": "richards@gmail.com",
    "time":"04:00 AM"
  }
])

期望的outcol查询结果

{"full_name": "Joel K", "email_address": "joelk@gmail.com", "collection_name":"col1"}
{"full_name": "Sam K", "email_address": "samk@gmail.com", "collection_name":"col1"}
{"full_name": "Karan H", "email_address": "karanh@gmail.com", "collection_name":"col2"}
{"full_name": "John H", "email_address": "john@gmail.com", "collection_name":"col2"}
{"full_name": "Tristan J", "email_address": "tristanj@gmail.com", "collection_name":"col3"}
{"full_name": "Richard S", "email_address": "richards@gmail.com", "collection_name":"col3"}

实现方法

可以用MongoDB的$unionWith聚合操作符合并多个集合,配合$project筛选公共字段并添加源集合标识,最后用$out将结果写入新集合outcol。

具体聚合命令

db.col1.aggregate([
  // 筛选公共字段,添加col1的来源标识
  {
    $project: {
      _id: 0,
      full_name: 1,
      email_address: 1,
      collection_name: "col1"
    }
  },
  // 合并col2的处理后数据
  {
    $unionWith: {
      coll: "col2",
      pipeline: [
        {
          $project: {
            _id: 0,
            full_name: 1,
            email_address: 1,
            collection_name: "col2"
          }
        }
      ]
    }
  },
  // 合并col3的处理后数据
  {
    $unionWith: {
      coll: "col3",
      pipeline: [
        {
          $project: {
            _id: 0,
            full_name: 1,
            email_address: 1,
            collection_name: "col3"
          }
        }
      ]
    }
  },
  // 将最终结果写入outcol集合
  {
    $out: "outcol"
  }
])

关键说明

  1. $project阶段:只保留需要的公共字段,排除默认的_id,同时添加collection_name标记数据来源。
  2. $unionWith阶段:依次合并其他集合的处理后数据,每个集合都用相同的字段筛选逻辑。
  3. $out阶段:将合并结果写入outcol,如果集合已存在会被覆盖;若需要追加数据,可替换为$merge操作符。

执行完上述命令后,查询db.outcol.find()即可得到期望结果。


内容的提问来源于stack exchange,提问作者Sid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 09:43:34