You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame转换为包含嵌套重复组结构的JSON文件

如何将Pandas DataFrame转换为包含嵌套重复组结构的JSON文件

我明白你遇到的问题了——你想把重复的客户行整合起来,每个客户只出现一次,同时把他们的所有保单详情放在一个数组里。之前的方法因为没有做分组聚合,导致每个保单都生成了一条重复的客户记录,而且保单详情也没有被整理成数组。咱们来一步步解决这个问题:

解决方案思路

核心是先按type和customer_id对DataFrame进行分组,然后对每个客户组做以下处理:

  1. 提取客户的通用信息(email、# of policies),同一客户的这些值完全重复,取第一行即可
  2. 将该客户的所有保单详情整理成一个数组
  3. 把这些信息组合成你需要的嵌套结构,最后转换为JSON格式

完整代码实现

import pandas as pd
import json

# 你的原始DataFrame
df = pd.DataFrame({
    'type': ['customer']*15,
    'customer_id': ['1-0000001']*4 + ['1-0000002']*6 + ['1-0000003']*5,
    'email': ['customer1@otenet.gr']*4 + ['customer2@gmail.com']*6 + ['customer3@yahoo.com.au']*5,
    '# of policies':[4]*4 + [6]*6 + [5]*5,
    'POLICY_NO': [f"00000000{i}" for i in range(1,16)],
    'RECEIPT_NO': [f"42000000{i}" for i in range(1,16)],
    'PAYMENT_CODE': [f"RF3500000000000000000000{i}" if i not in [7,10,13] else 'null' for i in range(1,16)],
    'KLADOS': ['Αυτοκινήτου']*15
})

# 初始化结果列表,用来存放每个客户的完整结构
result = []

# 按type和customer_id分组遍历每个客户的数据
for (customer_type, cust_id), group in df.groupby(['type', 'customer_id']):
    # 构建客户属性字典,包含通用信息和保单详情数组
    customer_attributes = {
        'email': group['email'].iloc[0],  # 取第一行的email(同一客户所有行都相同)
        '# of policies': group['# of policies'].iloc[0],  # 同理取保单数量
        # 把当前客户的所有保单详情转成字典列表
        'policies details': group[['POLICY_NO', 'RECEIPT_NO', 'PAYMENT_CODE', 'KLADOS']].to_dict(orient='records')
    }
    
    # 组装单个客户的完整JSON结构
    customer_entry = {
        'type': customer_type,
        'customer_id': cust_id,
        'attributes': customer_attributes
    }
    
    result.append(customer_entry)

# 转换为JSON字符串,保留希腊字符并格式化缩进
json_output = json.dumps(result, ensure_ascii=False, indent=4)
print(json_output)

# 可选:保存到JSON文件
with open('customer_policies.json', 'w', encoding='utf-8') as f:
    json.dump(result, f, ensure_ascii=False, indent=4)

预期输出效果

输出的JSON会完全符合你想要的结构,每个客户只出现一次,保单详情以数组形式嵌套:

[
    {
        "type": "customer",
        "customer_id": "1-0000001",
        "attributes": {
            "email": "customer1@otenet.gr",
            "# of policies": 4,
            "policies details": [
                {
                    "POLICY_NO": "000000001",
                    "RECEIPT_NO": "420000001",
                    "PAYMENT_CODE": "RF35000000000000000000001",
                    "KLADOS": "Αυτοκινήτου"
                },
                {
                    "POLICY_NO": "000000002",
                    "RECEIPT_NO": "420000002",
                    "PAYMENT_CODE": "RF35000000000000000000002",
                    "KLADOS": "Αυτοκινήτου"
                },
                // 该客户剩余2条保单详情
            ]
        }
    },
    // 客户2、客户3的条目结构类似
]

为什么之前的方法有问题?

  1. 缺少分组聚合:你之前直接用to_json(orient='records'),会把DataFrame的每一行都转成独立的JSON对象,所以每个保单都对应一条重复的客户记录。
  2. 保单详情未整理为数组:你给每行单独生成了policy_details字典,而不是把同一客户的所有保单详情收集成一个列表,导致每个条目里的policy_details是单个字典而非数组。

备注:内容来源于stack exchange,提问作者Alessandro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 18:48:08