You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pymongo和pandas将MongoDB嵌套数组文档转为指定结构DataFrame

实现步骤

完整实现代码如下,兼容factors为NaN的场景:

import pymongo
import pandas as pd

# 1. 连接MongoDB,按需替换连接地址、库名、集合名
client = pymongo.MongoClient("mongodb://localhost:27017/")
db = client["your_database_name"]
collection = db["your_collection_name"]

# 2. 查询全量数据,仅取需要的字段降低内存占用
raw_data = list(collection.find({}, {"name": 1, "factors": 1, "_id": 0}))

# 3. 展开嵌套数组,关联name字段
result_list = []
for doc in raw_data:
    current_name = doc.get("name")
    factors = doc.get("factors")
    # 处理factors为NaN的异常情况
    if pd.isna(factors):
        # 不需要保留NaN行可以直接写continue跳过该文档
        result_list.append({
            "name": current_name,
            "factorId": None,
            "Index": None,
            "weight": None
        })
        continue
    # 过滤非数组格式的异常factors值
    if not isinstance(factors, list):
        continue
    # 遍历数组生成每行数据
    for factor_item in factors:
        result_list.append({
            "name": current_name,
            "factorId": factor_item.get("factorId"),
            "Index": factor_item.get("Index"),
            "weight": factor_item.get("weight")
        })

# 4. 转换为DataFrame
df = pd.DataFrame(result_list)

# 如果你需要和示例一样把weight统一改为0,添加下面这行即可
# df["weight"] = 0

补充说明

  • 如果不需要保留factors为NaN的文档,在查询阶段就可以过滤,修改第二步的查询语句为:
raw_data = list(collection.find(
    {"factors": {"$not": {"$type": "double"}}}, # MongoDB中NaN存储为double类型,直接过滤
    {"name": 1, "factors": 1, "_id": 0}
))
  • 代码里用get方法取值是为了兼容个别文档缺失字段的情况,避免运行报错。

内容的提问来源于stack exchange,提问作者Kalindu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 03:57:03