You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将CSV转换为指定嵌套JSON结构(层级修正)

问题描述

我有一个包含customerNumber和itemNumber两列的CSV文件,需要将其转换为如下指定的嵌套JSON格式:

目标JSON示例:

[
    {
        "customerInformation": {
            "customerNumber": "000202"
        },
        "itemInformation": [
            {
                "itemNumber": "1660100110857"
            },
            {
                "itemNumber": "510520144"
            }
        ]
    },
    {
        "customerInformation": {
            "customerNumber": "001040"
        },
        "itemInformation": [
            {
                "itemNumber": "0171100243"
            },
            {
                "itemNumber": "0171100288"
            },
            {
                "itemNumber": "017110021212"
            },
            {
                "itemNumber": "017110561010"
            },
            {
                "itemNumber": "2028G1006"
            },
            {
                "itemNumber": "2028G206"
            },
            {
                "itemNumber": "0669406015"
            },
            {
                "itemNumber": "181902000"
            },
            {
                "itemNumber": "0669401020"
            }
        ]
    }
]

我编写的Pandas代码无法实现customerNumber嵌套在customerInformation下的结构,现有代码如下:

import pandas as pd
import os
import sys
import json

directory = os.path.join(os.path.join(os.environ['USERPROFILE']), 'Desktop')

dtype_dic = {'customerNumber' : str, 'itemNumber':str}

if os.path.isfile(directory + '/' + 'Price_Checker.csv'):
    print("Program Running.... please standby")
else:
    print("Required file Price_Checker.csv not found on user's Desktop")
    sys.exit()

df = pd.read_csv(directory + r'\\Price_Checker.csv', dtype=dtype_dic)


print(df)
result = []

for cust_id, cust_df in df.groupby('customerNumber'):
    cust_dict = {'customerInformation': cust_id}
    cust_dict['itemInformation'] = []
    for item_id, item_df in cust_df.groupby('itemNumber'):
        item_dict = {'itemNumber': item_id}
        cust_dict['itemInformation'].append(item_dict)
    result.append(cust_dict)

print(json.dumps(result, indent=4))
解决方案

你的代码核心问题在于构造customerInformation时直接传入了cust_id字符串,而目标格式要求它是一个包含customerNumber键的嵌套字典。另外,遍历itemNumber时无需额外groupby,直接提取去重后的列表即可简化代码。

修正后的完整代码:

import pandas as pd
import os
import sys
import json

directory = os.path.join(os.path.join(os.environ['USERPROFILE']), 'Desktop')

dtype_dic = {'customerNumber' : str, 'itemNumber':str}

if os.path.isfile(os.path.join(directory, 'Price_Checker.csv')):
    print("Program Running.... please standby")
else:
    print("Required file Price_Checker.csv not found on user's Desktop")
    sys.exit()

df = pd.read_csv(os.path.join(directory, 'Price_Checker.csv'), dtype=dtype_dic)

result = []

# 按customerNumber分组
for cust_id, cust_df in df.groupby('customerNumber'):
    # 构造正确的customerInformation嵌套结构
    cust_dict = {
        'customerInformation': {'customerNumber': cust_id},
        'itemInformation': []
    }
    # 提取当前客户的所有itemNumber,去重后构造字典列表
    for item_id in cust_df['itemNumber'].unique():
        cust_dict['itemInformation'].append({'itemNumber': item_id})
    result.append(cust_dict)

# 输出格式化后的JSON
print(json.dumps(result, indent=4))

关键修改点

  • 把{'customerInformation': cust_id}改为{'customerInformation': {'customerNumber': cust_id}},满足嵌套对象的格式要求。
  • 替换cust_df.groupby('itemNumber')为cust_df['itemNumber'].unique(),避免不必要的分组操作,直接获取当前客户的所有唯一商品编号。
  • 用os.path.join拼接文件路径,替代字符串拼接,提升代码跨平台兼容性。

内容的提问来源于stack exchange,提问作者Clare

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 22:52:39