如何在Python中为含嵌套文档列表的MongoDB文档创建聚合查询
First, let's recap your document structure for clarity:
{ "_id" : "132743", "RECORD_DATA" : [ { "FIELD_TYPE" : "Primary", "DATA" : "Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec blandit leo sit amet nisi ultricies bibendum. Aenean efficitur pharetra diam, non pretium nisi blandit eu. Maecenas eget dolor sed ipsum semper posuere id eget purus. Ut tempor massa vel porta euismod. Vivamus et elementum justo. Aliquam porta, ipsum at semper pulvinar, turpis ipsum congue orci, a fri..." } ] }
When dealing with nested arrays like RECORD_DATA, the MongoDB aggregation framework has several operators to help you query, transform, and analyze the data. Below are common use cases with Python code examples using pymongo:
1. Unwind the Array to Process Individual Elements
If you need to work with each object in the RECORD_DATA array as a separate document, use the $unwind stage. This is useful for filtering, aggregating, or transforming individual array elements.
Example: Extract all "Primary" type records with their parent document ID
from pymongo import MongoClient # Connect to MongoDB (adjust your connection string as needed) client = MongoClient("mongodb://localhost:27017/") db = client["your_database_name"] collection = db["your_collection_name"] # Define the aggregation pipeline pipeline = [ # Split each element in RECORD_DATA into its own document {"$unwind": "$RECORD_DATA"}, # Filter only entries where FIELD_TYPE is "Primary" {"$match": {"RECORD_DATA.FIELD_TYPE": "Primary"}}, # Keep only the fields we need in the output {"$project": { "_id": 1, "primary_content": "$RECORD_DATA.DATA" }} ] # Run the aggregation and convert results to a list results = list(collection.aggregate(pipeline)) # Print the output for doc in results: print(doc)
2. Filter Array Elements Without Unwinding
If you want to keep the parent document intact but only retain specific elements in the RECORD_DATA array, use $filter in a $project stage.
Example: Keep only "Primary" type entries in the RECORD_DATA array
pipeline = [ {"$project": { "RECORD_DATA": { "$filter": { "input": "$RECORD_DATA", "as": "item", "cond": {"$eq": ["$$item.FIELD_TYPE", "Primary"]} } } }} ] results = list(collection.aggregate(pipeline))
3. Aggregate Statistics Across Array Elements
You can calculate metrics like count, average length of DATA fields, etc., using aggregation stages.
Example: Count how many "Primary" records exist across all documents
pipeline = [ {"$unwind": "$RECORD_DATA"}, {"$match": {"RECORD_DATA.FIELD_TYPE": "Primary"}}, {"$group": { "_id": None, "total_primary_records": {"$sum": 1} }} ] results = list(collection.aggregate(pipeline)) print(f"Total Primary records found: {results[0]['total_primary_records']}")
4. Extract Specific Fields from Array Elements
If you need to pluck a specific field (like DATA) from all matching array elements into a new array, combine $map with $filter.
Example: Create an array of all "Primary" DATA values for each document
pipeline = [ {"$project": { "primary_data_list": { "$map": { "input": {"$filter": { "input": "$RECORD_DATA", "as": "item", "cond": {"$eq": ["$$item.FIELD_TYPE", "Primary"]} }}, "as": "filtered_item", "in": "$$filtered_item.DATA" } } }} ] results = list(collection.aggregate(pipeline))
Quick Tips for Better Performance:
- Place
$matchstages as early as possible in your pipeline to filter out unnecessary documents before processing. - If your
RECORD_DATAarray might be empty or missing, addpreserveNullAndEmptyArrays: trueto the$unwindstage to retain those documents. - Combine operators like
$sort,$limit, or$lookupfor more complex workflows as needed.
内容的提问来源于stack exchange,提问作者Deepak Aggarwal

