求助:高效转换不规则多层嵌套字典为Pandas DataFrame
高效转换多层嵌套字典为指定Pandas DataFrame
对于这种多层嵌套且结构不规则的字典,最高效的方式是通过生成器表达式扁平化数据,再直接转换为DataFrame——这种方式无需预先构建完整的列表,内存占用极低,非常适合处理超大体积的字典数据。
实现代码
import pandas as pd # 你的输入字典 a = {'Cat0': {'brand1': {'b': 0.78, 'c': 1}, 'brand2': {'k': 1, 'c': 1}}, 'Cat1': {'brand4': {'b': 10, 's': 0.0}}, 'Cat2': {'brand1': {'j': 1, 'c': 0.0}}} # 生成器:逐层级遍历并生成每行数据 flat_generator = ( (category, brand, peer, value) for category, brand_dict in a.items() for brand, peer_dict in brand_dict.items() for peer, value in peer_dict.items() ) # 转换为目标格式的DataFrame df = pd.DataFrame(flat_generator, columns=['Category', 'Brand', 'Peer', 'Value'])
效果验证
运行上述代码后,输出的DataFrame将完全匹配你指定的格式:
Category Brand Peer Value 0 Cat0 brand1 b 0.78 1 Cat0 brand1 c 1.00 2 Cat0 brand2 k 1.00 3 Cat0 brand2 c 1.00 4 Cat1 brand4 b 10.00 5 Cat1 brand4 s 0.00 6 Cat2 brand1 j 1.00 7 Cat2 brand1 c 0.00
为什么这种方法高效?
- 内存友好:生成器不会一次性将所有数据加载到内存中,而是在Pandas构建DataFrame时按需生成每一行,避免处理超大字典时的内存溢出问题。
- 遍历效率高:直接通过三层嵌套的生成器遍历字典,比先构建列表再转换的方式少了一次内存拷贝操作,整体执行速度更快。
内容的提问来源于stack exchange,提问作者Forinstance
相关产品推荐
相关产品推荐

