MongoDB中获取数组类型字段的去重值并写入目标集合的实现方案
Absolutely, you can solve this problem with an aggregation pipeline built to handle large datasets without hitting the 16MB cap from the distinct command. Let’s walk through the approach step by step:
Why your initial attempts hit limits
The distinct command has a hard 16MB limit on result size, which makes it unsuitable for your large arrays. Your aggregate attempt was on the right track—you just need to adjust the stages to collect and deduplicate values properly.
The Aggregation Pipeline
This pipeline will unwind the array, deduplicate values, collect them into a single array, and write the result to your target collection:
db.src_coll.aggregate([ // Step 1: Unwind the letters array into individual documents { $unwind: "$letters" }, // Step 2: Group all documents into one, collecting unique letters with $addToSet { $group: { _id: null, // Group all documents together all_letters: { $addToSet: "$letters" } // Auto-deduplicates as it collects } }, // Step 3: Write the result to the target collection { $out: "dst_coll" } ], { allowDiskUse: true }) // Critical for large datasets to avoid memory limits
Breaking down each stage:
- $unwind: Takes each element in the
lettersarray and creates a separate document for it. This lets us work with individual values instead of entire arrays. - $group: By setting
_id: null, we combine all unwound documents into a single group. The$addToSetoperator automatically adds only unique values to theall_lettersarray, so duplicates are removed automatically. - $out: Writes the final aggregated document directly to your
dst_collcollection. If the collection already exists, this will replace it (use$mergeinstead if you need to append to an existing collection). - allowallowDiskUse: true: Enables MongoDB to use disk storage for intermediate results—essential given your 314MB dataset—to prevent memory overflow errors.
Optional: Sort the result array
If you want the all_letters array to be sorted (matching your example), add these extra stages right before $out:
// Insert these after the first $group stage { $unwind: "$all_letters" }, { $sort: { "all_letters": 1 } }, { $group: { _id: null, all_letters: { $push: "$all_letters" } } }
This unwinds the deduplicated array, sorts individual elements, then re-groups them into a sorted array.
Verifying the result
After running the pipeline, check the target collection with:
db.dst_coll.find()
You’ll get a single document with the full list of unique letters values, just like your desired output.
内容的提问来源于stack exchange,提问作者Faaiz

