使用mongoexport实现MongoDB distinct查询结果导出的问题求助
distinct Query Results with mongoexport Let's break down your problem and fix this step by step.
First, your original mongoexport command has a few critical issues that are causing the error:
- You split the query and field selection into two separate objects, which isn't valid syntax for the
--queryparameter. The--queryexpects a single JSON object for filtering documents. - Your
post_idvalues are strings in the documents (e.g.,"28011986676_10155780942281677"), but you used a numeric value in{$gte:28011986676}—this won't match correctly since MongoDB treats strings and numbers as distinct types. - To specify which fields to export, you need to use the
--fieldsparameter, not append fields directly after the query.
Now, to get the exact deduplicated results as your db.sInsert.distinct() query, you have two solid options:
Option 1: Use the mongo Shell Directly (Simplest Exact Match)
If you just need to export the exact array output from your distinct query, run the command directly in the mongo shell and redirect the output to a file. This is the most straightforward way to replicate your original result:
mongo sDB --eval 'printjson(db.sInsert.distinct("comments_data.comments.data.message", {post_id: "28011986676"}))' > test.json
This will export the exact [ "Who else loves Apple ?" ] (or your actual deduplicated messages) as a JSON array into test.json, just like your original query returns.
Option 2: Use mongoexport with an Aggregation Pipeline
If you prefer using mongoexport (e.g., for more complex transformations), use an aggregation pipeline to replicate the distinct behavior. The pipeline filters documents, flattens nested arrays, deduplicates messages, and formats the output:
Aggregation Pipeline Breakdown:
$match: Filters documents to only those matching yourpost_idcondition (use a string value to match your document structure).$unwind: Flattens the nestedcomments_data.comments.dataarray to process individual messages.$group: Groups documents by themessagefield to remove duplicates (using_idas the message value).$project: Restructures output to show a cleanmessagefield instead of the_idused for grouping.- Optional
$sort: Sorts results alphabetically for readability.
Corresponding mongoexport Command:
mongoexport --db sDB --collection sInsert --aggregate '[ {"$match": {"post_id": "28011986676"}}, {"$unwind": "$comments_data.comments.data"}, {"$group": {"_id": "$comments_data.comments.data.message"}}, {"$project": {"message": "$_id", "_id": 0}}, {"$sort": {"message": 1}} ]' --out test.json
If you want the output to be a single array (like the distinct result) instead of individual objects, add a final $group stage to collect all messages into an array:
mongoexport --db sDB --collection sInsert --aggregate '[ {"$match": {"post_id": "28011986676"}}, {"$unwind": "$comments_data.comments.data"}, {"$group": {"_id": "$comments_data.comments.data.message"}}, {"$group": {"_id": null, "messages": {"$push": "$_id"}}}, {"$project": {"_id": 0}} ]' --out test.json
This will export a single object with a messages array containing all deduplicated comments, matching the spirit of your original distinct query.
内容的提问来源于stack exchange,提问作者xxxSL

