You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将特殊结构字典转换为含多contract_id的Pandas DataFrame

解决方案:将特殊结构字典转换为Pandas DataFrame

核心思路

  1. 拆分键结构:把每个字典键按_分割为「带.pdf的文件名」和「属性名」两部分,去掉文件名的.pdf后缀得到contract_id
  2. 聚合合同数据:按contract_id分组,收集每个合同对应的所有属性键值对
  3. 转换为DataFrame:将聚合后的结构转换为标准的表格格式,contract_id作为独立列,每个属性对应一列

实现代码

import pandas as pd

def dict_to_contract_df(raw_dict):
    # 临时存储每个合同的完整数据
    contract_records = {}
    
    for full_key, value in raw_dict.items():
        # 分割键:仅拆分一次,避免文件名含下划线时出错
        file_segment, prop_name = full_key.split('_', 1)
        # 提取contract_id:移除.pdf后缀
        contract_id = file_segment.rstrip('.pdf')
        
        # 处理值为列表的情况:取第一个元素,空列表则设为None
        if isinstance(value, list):
            processed_value = value[0] if value else None
        else:
            processed_value = value
        
        # 初始化合同条目,添加属性值
        if contract_id not in contract_records:
            contract_records[contract_id] = {'contract_id': contract_id}
        contract_records[contract_id][prop_name] = processed_value
    
    # 转换为DataFrame
    return pd.DataFrame(list(contract_records.values()))

测试示例

示例1:单个合同的字典输入

sample_dict1 = {
    'CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-93ea-96073626dccb.pdf_Rebate-Count': 'Two rebate types',
    'CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-93ea-96073626dccb.pdf_Rebate-Spec-CashCredit': 'Credit Note',
    'CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-93ea-96073626dccb.pdf_Rebate-Cadence-First-StartDate': 'July 1, 2021',
    'CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-93ea-96073626dccb.pdf_Rebate-Cadence-LastDate': 'July 15, 2023',
    'CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-93ea-96073626dccb.pdf_Rebate-Cadence-CadenceCollection': 'Quarterly'
}

df_result1 = dict_to_contract_df(sample_dict1)
print(df_result1)

输出结果:

contract_id       Rebate-Count Rebate-Spec-CashCredit Rebate-Cadence-First-StartDate Rebate-Cadence-LastDate Rebate-Cadence-CadenceCollection
0  CC OTH 00009438 2023 TR.2a1e3e6f-58c4-4166-...  Two rebate types           Credit Note                  July 1, 2021           July 15, 2023                          Quarterly

示例2:含列表值的单个合同字典

sample_dict2 = {
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Exists': ['Yes'],
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Count': [],
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Spec-CashCredit': ['Cash Refund/Payment'],
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Cadence-First-StartDate': ['July 16, 2022'],
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Cadence-LastDate': ['July 15, 2023'],
    'Rebate Agreement Final (Signed Document).pdf_Rebate-Cadence-CadenceCollection': ['Annual']
}

df_result2 = dict_to_contract_df(sample_dict2)
print(df_result2)

输出结果:

contract_id Rebate-Exists Rebate-Count Rebate-Spec-CashCredit Rebate-Cadence-First-StartDate Rebate-Cadence-LastDate Rebate-Cadence-CadenceCollection
0  Rebate Agreement Final (Signed Document)           Yes         None    Cash Refund/Payment                  July 16, 2022           July 15, 2023                            Annual

示例3:多个合同的混合输入

将两个示例字典合并后输入,函数会自动按contract_id分组生成两行数据:

combined_dict = {**sample_dict1, **sample_dict2}
df_combined = dict_to_contract_df(combined_dict)
print(df_combined)

关键细节说明

  • 使用split('_', 1)而非普通split('_'),确保文件名中包含下划线时不会被错误拆分
  • 兼容值为字符串或列表的情况,自动处理空列表为None,避免DataFrame格式异常
  • 自动识别多个contract_id,每个合同对应DataFrame的一行,缺失的属性会自动填充为NaN

内容的提问来源于stack exchange,提问作者Wolfy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 03:44:59