如何将Pandas DataFrame转换为包含嵌套重复组结构的JSON文件
如何将Pandas DataFrame转换为包含嵌套重复组结构的JSON文件
我明白你遇到的问题了——你想把重复的客户行整合起来,每个客户只出现一次,同时把他们的所有保单详情放在一个数组里。之前的方法因为没有做分组聚合,导致每个保单都生成了一条重复的客户记录,而且保单详情也没有被整理成数组。咱们来一步步解决这个问题:
解决方案思路
核心是先按type和customer_id对DataFrame进行分组,然后对每个客户组做以下处理:
- 提取客户的通用信息(
email、# of policies),同一客户的这些值完全重复,取第一行即可 - 将该客户的所有保单详情整理成一个数组
- 把这些信息组合成你需要的嵌套结构,最后转换为JSON格式
完整代码实现
import pandas as pd import json # 你的原始DataFrame df = pd.DataFrame({ 'type': ['customer']*15, 'customer_id': ['1-0000001']*4 + ['1-0000002']*6 + ['1-0000003']*5, 'email': ['customer1@otenet.gr']*4 + ['customer2@gmail.com']*6 + ['customer3@yahoo.com.au']*5, '# of policies':[4]*4 + [6]*6 + [5]*5, 'POLICY_NO': [f"00000000{i}" for i in range(1,16)], 'RECEIPT_NO': [f"42000000{i}" for i in range(1,16)], 'PAYMENT_CODE': [f"RF3500000000000000000000{i}" if i not in [7,10,13] else 'null' for i in range(1,16)], 'KLADOS': ['Αυτοκινήτου']*15 }) # 初始化结果列表,用来存放每个客户的完整结构 result = [] # 按type和customer_id分组遍历每个客户的数据 for (customer_type, cust_id), group in df.groupby(['type', 'customer_id']): # 构建客户属性字典,包含通用信息和保单详情数组 customer_attributes = { 'email': group['email'].iloc[0], # 取第一行的email(同一客户所有行都相同) '# of policies': group['# of policies'].iloc[0], # 同理取保单数量 # 把当前客户的所有保单详情转成字典列表 'policies details': group[['POLICY_NO', 'RECEIPT_NO', 'PAYMENT_CODE', 'KLADOS']].to_dict(orient='records') } # 组装单个客户的完整JSON结构 customer_entry = { 'type': customer_type, 'customer_id': cust_id, 'attributes': customer_attributes } result.append(customer_entry) # 转换为JSON字符串,保留希腊字符并格式化缩进 json_output = json.dumps(result, ensure_ascii=False, indent=4) print(json_output) # 可选:保存到JSON文件 with open('customer_policies.json', 'w', encoding='utf-8') as f: json.dump(result, f, ensure_ascii=False, indent=4)
预期输出效果
输出的JSON会完全符合你想要的结构,每个客户只出现一次,保单详情以数组形式嵌套:
[ { "type": "customer", "customer_id": "1-0000001", "attributes": { "email": "customer1@otenet.gr", "# of policies": 4, "policies details": [ { "POLICY_NO": "000000001", "RECEIPT_NO": "420000001", "PAYMENT_CODE": "RF35000000000000000000001", "KLADOS": "Αυτοκινήτου" }, { "POLICY_NO": "000000002", "RECEIPT_NO": "420000002", "PAYMENT_CODE": "RF35000000000000000000002", "KLADOS": "Αυτοκινήτου" }, // 该客户剩余2条保单详情 ] } }, // 客户2、客户3的条目结构类似 ]
为什么之前的方法有问题?
- 缺少分组聚合:你之前直接用
to_json(orient='records'),会把DataFrame的每一行都转成独立的JSON对象,所以每个保单都对应一条重复的客户记录。 - 保单详情未整理为数组:你给每行单独生成了
policy_details字典,而不是把同一客户的所有保单详情收集成一个列表,导致每个条目里的policy_details是单个字典而非数组。
备注:内容来源于stack exchange,提问作者Alessandro
相关产品推荐
相关产品推荐

