You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB中获取数组类型字段的去重值并写入目标集合的实现方案

Solution: Deduplicate Array Field and Write to New Collection

Absolutely, you can solve this problem with an aggregation pipeline built to handle large datasets without hitting the 16MB cap from the distinct command. Let’s walk through the approach step by step:

Why your initial attempts hit limits

The distinct command has a hard 16MB limit on result size, which makes it unsuitable for your large arrays. Your aggregate attempt was on the right track—you just need to adjust the stages to collect and deduplicate values properly.

The Aggregation Pipeline

This pipeline will unwind the array, deduplicate values, collect them into a single array, and write the result to your target collection:

db.src_coll.aggregate([
  // Step 1: Unwind the letters array into individual documents
  { $unwind: "$letters" },
  // Step 2: Group all documents into one, collecting unique letters with $addToSet
  {
    $group: {
      _id: null, // Group all documents together
      all_letters: { $addToSet: "$letters" } // Auto-deduplicates as it collects
    }
  },
  // Step 3: Write the result to the target collection
  { $out: "dst_coll" }
], { allowDiskUse: true }) // Critical for large datasets to avoid memory limits

Breaking down each stage:

  • $unwind: Takes each element in the letters array and creates a separate document for it. This lets us work with individual values instead of entire arrays.
  • $group: By setting _id: null, we combine all unwound documents into a single group. The $addToSet operator automatically adds only unique values to the all_letters array, so duplicates are removed automatically.
  • $out: Writes the final aggregated document directly to your dst_coll collection. If the collection already exists, this will replace it (use $merge instead if you need to append to an existing collection).
  • allowallowDiskUse: true: Enables MongoDB to use disk storage for intermediate results—essential given your 314MB dataset—to prevent memory overflow errors.

Optional: Sort the result array

If you want the all_letters array to be sorted (matching your example), add these extra stages right before $out:

// Insert these after the first $group stage
{ $unwind: "$all_letters" },
{ $sort: { "all_letters": 1 } },
{
  $group: {
    _id: null,
    all_letters: { $push: "$all_letters" }
  }
}

This unwinds the deduplicated array, sorts individual elements, then re-groups them into a sorted array.

Verifying the result

After running the pipeline, check the target collection with:

db.dst_coll.find()

You’ll get a single document with the full list of unique letters values, just like your desired output.

内容的提问来源于stack exchange,提问作者Faaiz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:42:34