You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将DataFrame转换为指定结构的JSON?

将DataFrame转换为指定嵌套结构JSON的解决方案

目标JSON包含顶层单值字段、嵌套数组字段products和列表字段documentationRequestedName,属于混合嵌套结构,直接将其转为DataFrame会因为层级问题失败。以下是正确的转换步骤和代码:

1. 基础场景:已有产品DataFrame和顶层字段值

如果已经拆分好产品数据和顶层字段数据,直接组合后转JSON即可:

import pandas as pd
import json

# 构造产品数据的DataFrame
products_df = pd.DataFrame([
    {"productId": 123456790, "quantaty": 10, "orderLine": "10", "buyerRef": "my ref"},
    {"productId": 123456791, "quantaty": 15, "orderLine": "20", "buyerRef": "my ref"}
])

# 定义顶层字段和文档列表
top_level_data = {
    "costId": 109,
    "paymentMethod": "termTransferWire",
    "totalPriceTaxIncl": 1200,
    "deliveryAddressId": 218,
    "buyerOrderNumber": "Test",
    "documentationRequestedName": ["document 1", "document 2"]
}

# 组合成目标结构字典
result_dict = {**top_level_data, "products": products_df.to_dict("records")}

# 转换为格式化后的JSON
target_json = json.dumps(result_dict, indent=4)
print(target_json)
  • 关键:products_df.to_dict("records")会把DataFrame每行转为一个字典,正好匹配products的数组结构。

2. 进阶场景:原始DataFrame包含所有重复的顶层字段

如果你的原始DataFrame是扁平结构(每行产品数据重复携带顶层字段),可以这样提取转换:

import pandas as pd
import json

# 示例原始扁平DataFrame
raw_df = pd.DataFrame([
    {"productId": 123456790, "quantaty": 10, "orderLine": "10", "buyerRef": "my ref",
     "costId": 109, "paymentMethod": "termTransferWire", "totalPriceTaxIncl": 1200,
     "deliveryAddressId": 218, "buyerOrderNumber": "Test", "doc1": "document 1", "doc2": "document 2"},
    {"productId": 123456791, "quantaty": 15, "orderLine": "20", "buyerRef": "my ref",
     "costId": 109, "paymentMethod": "termTransferWire", "totalPriceTaxIncl": 1200,
     "deliveryAddressId": 218, "buyerOrderNumber": "Test", "doc1": "document 1", "doc2": "document 2"}
])

# 提取顶层字段(取第一行值即可,因为所有行重复)
top_level = {
    "costId": raw_df.iloc[0]["costId"],
    "paymentMethod": raw_df.iloc[0]["paymentMethod"],
    "totalPriceTaxIncl": raw_df.iloc[0]["totalPriceTaxIncl"],
    "deliveryAddressId": raw_df.iloc[0]["deliveryAddressId"],
    "buyerOrderNumber": raw_df.iloc[0]["buyerOrderNumber"],
    "documentationRequestedName": [raw_df.iloc[0]["doc1"], raw_df.iloc[0]["doc2"]]
}

# 提取产品数据(仅保留产品相关列)
products = raw_df[["productId", "quantaty", "orderLine", "buyerRef"]].to_dict("records")

# 组合并生成JSON
result_dict = {**top_level, "products": products}
target_json = json.dumps(result_dict, indent=4)
print(target_json)

内容的提问来源于stack exchange,提问作者user3833880

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 23:35:36