You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效将大型列表字典转换为Pandas DataFrame?

优化大规模字典转DataFrame的性能问题

问题背景

  • 输入为列表字典,所有列表长度一致,数据类型包含浮点数、长短字符串、大小整数、布尔值等,来自多个子系统
  • 传统pd.DataFrame.from_dict方法处理18M行×100列的数据集时,耗时超15分钟,性能瓶颈明显

已尝试的优化方案及耗时

方案1:通过pd.concat批量生成Series后合并

ptable = pd.concat({k: pd.Series(v) for k, v in tbl_dict.items()}, axis=1)
  • 耗时:7分钟

方案2:逐列构建DataFrame

temp = pd.DataFrame()
i = 0
s = pd.Series(tbl_dict)  # 此步骤瞬间完成
cols = list(s.index)

# 逐列添加数据构建表格
while i < len(cols):
    c = cols[i]
    temp[c] = s[i]
    i += 1
  • 耗时:5-6分钟(同时将列数从60扩展至100)

模拟数据示例(18M行×100列)

{
    'ID1' : ['8abcdefgi','9abcdefgi', '1abcdefgi',....],
    'ID2' : [11111,12345,10001, ....],
    'ID3' : [123456789012, 123456789015, 323450789012, ...],
    'ID4' : [553456483415, 444879893478, 893012334091, ...],
    'ID5' : [1233445845, 7843212314, 3843312344, ...],
    'Status' : ['Live', 'Not Live', 'Live',...],
    'Sector' : ['Finance', 'Finance', 'Finance',...], 
    'Name1' : ['Tdjolkajdshytjsbggh-klhn', 'Adgtlkajdyxykgsbgddg-utyj',...],
    'Name2' : ['Equities Desk', None, 'Equities Desk', 'Commodities Desk',...],
    'ID6' : [76534, 90654, 78456,...],
    'Name3' : ['ABC', 'BCD', None,...],
    'Type1' : [None, None, None,...],
    'Flag1' : [0,0,1,0,...],
    'Flag2' : ['true', 'false', 'true', 'true',...],
    'ID7' : [123456789, 873631826, 876391723, 316071258,...],
    'Type2' : ['ABC', 'DDD', 'AHY',...], 
    'Name4' : ['Desk Name', 'Desk Name', 'Desk Name',...], 
    ...  # 后续63列,80%为Flag2类型(字符串布尔值),其余为Type1类型(全空值)
}

需求

寻求进一步优化的方案,将18M行×100列的字典转DataFrame的耗时压缩至更短时间。

内容的提问来源于stack exchange,提问作者GaryChin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 22:17:17