如何用Pandas将CSV转换为指定嵌套JSON结构(层级修正)
问题描述
我有一个包含customerNumber和itemNumber两列的CSV文件,需要将其转换为如下指定的嵌套JSON格式:
目标JSON示例:
[ { "customerInformation": { "customerNumber": "000202" }, "itemInformation": [ { "itemNumber": "1660100110857" }, { "itemNumber": "510520144" } ] }, { "customerInformation": { "customerNumber": "001040" }, "itemInformation": [ { "itemNumber": "0171100243" }, { "itemNumber": "0171100288" }, { "itemNumber": "017110021212" }, { "itemNumber": "017110561010" }, { "itemNumber": "2028G1006" }, { "itemNumber": "2028G206" }, { "itemNumber": "0669406015" }, { "itemNumber": "181902000" }, { "itemNumber": "0669401020" } ] } ]
我编写的Pandas代码无法实现customerNumber嵌套在customerInformation下的结构,现有代码如下:
import pandas as pd import os import sys import json directory = os.path.join(os.path.join(os.environ['USERPROFILE']), 'Desktop') dtype_dic = {'customerNumber' : str, 'itemNumber':str} if os.path.isfile(directory + '/' + 'Price_Checker.csv'): print("Program Running.... please standby") else: print("Required file Price_Checker.csv not found on user's Desktop") sys.exit() df = pd.read_csv(directory + r'\\Price_Checker.csv', dtype=dtype_dic) print(df) result = [] for cust_id, cust_df in df.groupby('customerNumber'): cust_dict = {'customerInformation': cust_id} cust_dict['itemInformation'] = [] for item_id, item_df in cust_df.groupby('itemNumber'): item_dict = {'itemNumber': item_id} cust_dict['itemInformation'].append(item_dict) result.append(cust_dict) print(json.dumps(result, indent=4))
解决方案
你的代码核心问题在于构造customerInformation时直接传入了cust_id字符串,而目标格式要求它是一个包含customerNumber键的嵌套字典。另外,遍历itemNumber时无需额外groupby,直接提取去重后的列表即可简化代码。
修正后的完整代码:
import pandas as pd import os import sys import json directory = os.path.join(os.path.join(os.environ['USERPROFILE']), 'Desktop') dtype_dic = {'customerNumber' : str, 'itemNumber':str} if os.path.isfile(os.path.join(directory, 'Price_Checker.csv')): print("Program Running.... please standby") else: print("Required file Price_Checker.csv not found on user's Desktop") sys.exit() df = pd.read_csv(os.path.join(directory, 'Price_Checker.csv'), dtype=dtype_dic) result = [] # 按customerNumber分组 for cust_id, cust_df in df.groupby('customerNumber'): # 构造正确的customerInformation嵌套结构 cust_dict = { 'customerInformation': {'customerNumber': cust_id}, 'itemInformation': [] } # 提取当前客户的所有itemNumber,去重后构造字典列表 for item_id in cust_df['itemNumber'].unique(): cust_dict['itemInformation'].append({'itemNumber': item_id}) result.append(cust_dict) # 输出格式化后的JSON print(json.dumps(result, indent=4))
关键修改点
- 把
{'customerInformation': cust_id}改为{'customerInformation': {'customerNumber': cust_id}},满足嵌套对象的格式要求。 - 替换
cust_df.groupby('itemNumber')为cust_df['itemNumber'].unique(),避免不必要的分组操作,直接获取当前客户的所有唯一商品编号。 - 用
os.path.join拼接文件路径,替代字符串拼接,提升代码跨平台兼容性。
内容的提问来源于stack exchange,提问作者Clare
相关产品推荐
相关产品推荐

