You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas实现将字典列表列转换为独立列的技术问题

问题描述

原始数据

grp = ["A","B","C","A","C","C","B"]
dictl = ["[{'TypeID': 0, 'Description': 'blah', 'DateCreated': '2018-08-09T14:00:30.957'}]",
"[{'TypeID': 0, 'Description': 'blah', 'DateCreated': '2018-08-09T14:00:30.957'}]",
"[]","[{'TypeID': 0, 'Description': 'blah', 'DateCreated': '2018-08-09T14:00:31.504'}]",
"[{'TypeID': 0, 'Description': 'blah', 'DateCreated': '2018-08-09T14:00:31.504'}]",
"[]","[{'TypeID': 0, 'Description': 'blah', 'DateCreated': '2018-08-09T14:00:31.504'}]"]
df = pd.DataFrame({'grp':grp,'dictl':dictl})

目标格式

pd.DataFrame({'grp':["A","B","C","A","C","C","B"],
              'TypeID':["0","0","","0","0","","0"],
              'Description':["blah","blah","","blah","blah","","blah"],
              'DateCreated':["2018-08-09T14:00:30.957","2018-08-09T14:00:30.957","","2018-08-09T14:00:31.504","2018-08-09T14:00:31.504","","2018-08-09T14:00:31.504"]})

尝试过的方法及报错

  • 方法1:
for grp, dictl in df:
    rec = {'Name': grp}
    rec.update(x for d in dictl for x in d.items())
    records.append(rec)

报错:ValueError: too many values to unpack (expected 2)

  • 方法2:
df['dictl'].apply(lambda c:
                                  pd.Series({next(iter(x.keys())).strip(':'):
                                             next(iter(x.values())) for x in c})
                                  )

报错:AttributeError: 'str' object has no attribute 'keys'

核心需求:数据量超200万行,需高效处理方案。


高效解决方案

针对大数据量,优先采用矢量化操作替代循环或逐行apply,避免性能瓶颈,步骤如下:

  1. 解析字符串为原生列表字典
    使用ast.literal_eval将dictl列的字符串格式数据转换为Python原生的列表(含字典或空列表),这一步是矢量化处理,效率远高于逐行解析:

    import ast
    df['dictl_parsed'] = df['dictl'].apply(ast.literal_eval)
    
  2. 提取字典并展开为列
    针对每个解析后的列表,空列表返回空字典,非空列表取第一个字典(匹配你的数据结构),再将这些字典展开为独立列:

    def extract_dict(lst):
        return lst[0] if lst else {}
    
    # 批量提取并展开为DataFrame
    dict_df = df['dictl_parsed'].apply(extract_dict).apply(pd.Series)
    
  3. 合并数据并格式化
    将原始grp列与展开后的列合并,同时把空值替换为字符串空,并将TypeID转为字符串类型(匹配目标格式):

    result = pd.concat([df['grp'], dict_df], axis=1).fillna('')
    result['TypeID'] = result['TypeID'].astype(str)
    

最终result的结构和内容与目标格式完全一致,且整个流程针对百万级数据做了性能优化。

内容的提问来源于stack exchange,提问作者frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 03:26:00