You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame转换为分组嵌套JSON以插入MongoDB?

解决方案

可以用Pandas的内置方法快速实现这种嵌套结构转换,以下是两种高效的方案:

方法1:逐行构造嵌套字段(直观易读)

先构造示例DataFrame用于测试:

import pandas as pd

df = pd.DataFrame({
    'name': ['robert', 'helena', 'cleito'],
    'order': ['car', 'car', 'car'],
    'age': [25, 45, 35],
    'color': ['red', 'yellow', 'green'],
    'rank': [5, 4, 4.5]
})

通过apply为每行生成details嵌套数组,再移除原字段:

# 生成details列:将color和rank打包为数组内的字典
df['details'] = df.apply(lambda row: [{'color': row['color'], 'rank': row['rank']}], axis=1)
# 删除不需要的原字段
df = df.drop(['color', 'rank'], axis=1)
# 转为目标格式的JSON
result_json = df.to_json(orient='records', indent=2)
print(result_json)

输出的JSON结构完全匹配需求,可直接用于MongoDB插入。

方法2:向量化构造(大数据集更高效)

如果处理的数据集行数较多,逐行apply效率偏低,可改用向量化方式构造嵌套字段:

# 批量生成details列
df['details'] = [[{'color': c, 'rank': r}] for c, r in zip(df['color'], df['rank'])]
# 删除原字段并转JSON
df = df.drop(['color', 'rank'], axis=1)
result_json = df.to_json(orient='records', indent=2)

直接插入MongoDB的优化方案

无需先转JSON,可直接将处理后的DataFrame转为字典列表插入,效率更高:

from pymongo import MongoClient

# 连接MongoDB
client = MongoClient('mongodb://localhost:27017/')
db = client['your_database']
collection = db['your_collection']

# 转为字典列表
data = df.to_dict('records')
# 批量插入
collection.insert_many(data)

内容的提问来源于stack exchange,提问作者Robert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:55:20