You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中为含嵌套文档列表的MongoDB文档创建聚合查询

Working with Nested Document Arrays in MongoDB Aggregation (Python)

First, let's recap your document structure for clarity:

{ 
  "_id" : "132743", 
  "RECORD_DATA" : [ 
    { 
      "FIELD_TYPE" : "Primary", 
      "DATA" : "Lorem ipsum dolor sit amet, consectetur adipiscing elit. Donec blandit leo sit amet nisi ultricies bibendum. Aenean efficitur pharetra diam, non pretium nisi blandit eu. Maecenas eget dolor sed ipsum semper posuere id eget purus. Ut tempor massa vel porta euismod. Vivamus et elementum justo. Aliquam porta, ipsum at semper pulvinar, turpis ipsum congue orci, a fri..." 
    } 
  ] 
}

When dealing with nested arrays like RECORD_DATA, the MongoDB aggregation framework has several operators to help you query, transform, and analyze the data. Below are common use cases with Python code examples using pymongo:

1. Unwind the Array to Process Individual Elements

If you need to work with each object in the RECORD_DATA array as a separate document, use the $unwind stage. This is useful for filtering, aggregating, or transforming individual array elements.

Example: Extract all "Primary" type records with their parent document ID

from pymongo import MongoClient

# Connect to MongoDB (adjust your connection string as needed)
client = MongoClient("mongodb://localhost:27017/")
db = client["your_database_name"]
collection = db["your_collection_name"]

# Define the aggregation pipeline
pipeline = [
    # Split each element in RECORD_DATA into its own document
    {"$unwind": "$RECORD_DATA"},
    # Filter only entries where FIELD_TYPE is "Primary"
    {"$match": {"RECORD_DATA.FIELD_TYPE": "Primary"}},
    # Keep only the fields we need in the output
    {"$project": {
        "_id": 1,
        "primary_content": "$RECORD_DATA.DATA"
    }}
]

# Run the aggregation and convert results to a list
results = list(collection.aggregate(pipeline))

# Print the output
for doc in results:
    print(doc)

2. Filter Array Elements Without Unwinding

If you want to keep the parent document intact but only retain specific elements in the RECORD_DATA array, use $filter in a $project stage.

Example: Keep only "Primary" type entries in the RECORD_DATA array

pipeline = [
    {"$project": {
        "RECORD_DATA": {
            "$filter": {
                "input": "$RECORD_DATA",
                "as": "item",
                "cond": {"$eq": ["$$item.FIELD_TYPE", "Primary"]}
            }
        }
    }}
]

results = list(collection.aggregate(pipeline))

3. Aggregate Statistics Across Array Elements

You can calculate metrics like count, average length of DATA fields, etc., using aggregation stages.

Example: Count how many "Primary" records exist across all documents

pipeline = [
    {"$unwind": "$RECORD_DATA"},
    {"$match": {"RECORD_DATA.FIELD_TYPE": "Primary"}},
    {"$group": {
        "_id": None,
        "total_primary_records": {"$sum": 1}
    }}
]

results = list(collection.aggregate(pipeline))
print(f"Total Primary records found: {results[0]['total_primary_records']}")

4. Extract Specific Fields from Array Elements

If you need to pluck a specific field (like DATA) from all matching array elements into a new array, combine $map with $filter.

Example: Create an array of all "Primary" DATA values for each document

pipeline = [
    {"$project": {
        "primary_data_list": {
            "$map": {
                "input": {"$filter": {
                    "input": "$RECORD_DATA",
                    "as": "item",
                    "cond": {"$eq": ["$$item.FIELD_TYPE", "Primary"]}
                }},
                "as": "filtered_item",
                "in": "$$filtered_item.DATA"
            }
        }
    }}
]

results = list(collection.aggregate(pipeline))

Quick Tips for Better Performance:

  • Place $match stages as early as possible in your pipeline to filter out unnecessary documents before processing.
  • If your RECORD_DATA array might be empty or missing, add preserveNullAndEmptyArrays: true to the $unwind stage to retain those documents.
  • Combine operators like $sort, $limit, or $lookup for more complex workflows as needed.

内容的提问来源于stack exchange,提问作者Deepak Aggarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:40:08