使用pymongo和pandas将MongoDB嵌套数组文档转为指定结构DataFrame
实现步骤
完整实现代码如下,兼容factors为NaN的场景:
import pymongo import pandas as pd # 1. 连接MongoDB,按需替换连接地址、库名、集合名 client = pymongo.MongoClient("mongodb://localhost:27017/") db = client["your_database_name"] collection = db["your_collection_name"] # 2. 查询全量数据,仅取需要的字段降低内存占用 raw_data = list(collection.find({}, {"name": 1, "factors": 1, "_id": 0})) # 3. 展开嵌套数组,关联name字段 result_list = [] for doc in raw_data: current_name = doc.get("name") factors = doc.get("factors") # 处理factors为NaN的异常情况 if pd.isna(factors): # 不需要保留NaN行可以直接写continue跳过该文档 result_list.append({ "name": current_name, "factorId": None, "Index": None, "weight": None }) continue # 过滤非数组格式的异常factors值 if not isinstance(factors, list): continue # 遍历数组生成每行数据 for factor_item in factors: result_list.append({ "name": current_name, "factorId": factor_item.get("factorId"), "Index": factor_item.get("Index"), "weight": factor_item.get("weight") }) # 4. 转换为DataFrame df = pd.DataFrame(result_list) # 如果你需要和示例一样把weight统一改为0,添加下面这行即可 # df["weight"] = 0
补充说明
- 如果不需要保留factors为NaN的文档,在查询阶段就可以过滤,修改第二步的查询语句为:
raw_data = list(collection.find( {"factors": {"$not": {"$type": "double"}}}, # MongoDB中NaN存储为double类型,直接过滤 {"name": 1, "factors": 1, "_id": 0} ))
- 代码里用
get方法取值是为了兼容个别文档缺失字段的情况,避免运行报错。
内容的提问来源于stack exchange,提问作者Kalindu
相关产品推荐
相关产品推荐

