You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB 4.2分片集群:三层嵌套数组指定字段提取(避$unwind)

MongoDB嵌套文档高效提取需求及补充方案

原始MongoDB文档结构:

{"_id":1,"_a":[{"_aId":{"CC":"CA"},"_p":[{"_pId":{"CC":"CA","SN":1},"s":{"c":{"t":[{"sId":1,"language":"CA","format":"A4"},{"sId":2,"language":"JP","format":"A4"}]},"a":{"t":[{"sId":4,"language":"CA","format":"A4"},{"sId":5,"language":"EN","format":"A3"}]},"d":{"t":[{"sId":10,"language":"CA","format":"A4"}]}}}]}]}

需求说明

需要高效提取满足以下条件的子文档字段:

  • _a._p._pId.CC 等于 "CA"
  • 子文档的 language 等于 "CA" 且 format 等于 "A4"

最终输出需包含 _aId、_pId、sId 三个字段,期望输出格式如下:

{_aId:{"CC": "CA"},_pId:{CC:"CA",SN:1},sId:1}
{_aId:{"CC": "CA"},_pId:{CC:"CA",SN:1},sId:4}
{_aId:{"CC": "CA"},_pId:{CC:"CA",SN:1},sId:10}

背景

此前已有类似问题的解决方案,该方案避免了高开销的$unwind操作,但未包含_aId字段,现需在原方案基础上补充该字段,环境为MongoDB分片集群4.2。

优化后的聚合管道方案

核心思路是在嵌套数组处理过程中保留_aId,并将所有符合条件的t数组条目扁平化,最终映射出目标字段:

完整聚合语句

db.collection.aggregate([
  {
    $project: {
      items: {
        $map: {
          input: "$_a",
          as: "aItem",
          in: {
            _aId: "$$aItem._aId",
            pItems: {
              $map: {
                input: {
                  $filter: {
                    input: "$$aItem._p",
                    cond: { $eq: ["$$this._pId.CC", "CA"] }
                  }
                },
                as: "pItem",
                in: {
                  _pId: "$$pItem._pId",
                  validEntries: {
                    $concatArrays: [
                      {
                        $filter: {
                          input: "$$pItem.s.c.t",
                          cond: { $and: [{ $eq: ["$$this.language", "CA"] }, { $eq: ["$$this.format", "A4"] }] }
                        }
                      },
                      {
                        $filter: {
                          input: "$$pItem.a.t",
                          cond: { $and: [{ $eq: ["$$this.language", "CA"] }, { $eq: ["$$this.format", "A4"] }] }
                        }
                      },
                      {
                        $filter: {
                          input: "$$pItem.d.t",
                          cond: { $and: [{ $eq: ["$$this.language", "CA"] }, { $eq: ["$$this.format", "A4"] }] }
                        }
                      }
                    ]
                  }
                }
              }
            }
          }
        }
      }
    }
  },
  {
    $project: {
      result: {
        $reduce: {
          input: "$items",
          initialValue: [],
          in: {
            $concatArrays: [
              "$$value",
              {
                $reduce: {
                  input: "$$this.pItems",
                  initialValue: [],
                  in: {
                    $concatArrays: [
                      "$$value",
                      {
                        $map: {
                          input: "$$this.validEntries",
                          as: "entry",
                          in: {
                            _aId: "$$this._aId",
                            _pId: "$$this._pId",
                            sId: "$$entry.sId"
                          }
                        }
                      }
                    ]
                  }
                }
              }
            ]
          }
        }
      }
    }
  },
  { $unwind: "$result" },
  { $replaceRoot: { newRoot: "$result" } }
])

逻辑说明

  1. 第一层映射与筛选:遍历_a数组,保留每个条目的_aId,同时筛选出_p数组中_pId.CC为"CA"的条目
  2. 合并有效子条目:将每个_p条目中s.c.t、a.t、d.t三个数组的符合条件条目合并为validEntries数组
  3. 扁平化结果集:通过两次$reduce嵌套,将多层嵌套的有效条目展开为一维数组,再通过$unwind和$replaceRoot输出最终格式的文档

内容的提问来源于stack exchange,提问作者R2D2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 23:24:37