如何用Pandas内置函数高效将四级字典转换为DataFrame?
问题:Pandas是否有内置函数无需循环转换嵌套字典为DataFrame?
我有一个记录不同国家冰淇淋销量的四级字典,具体结构如下:
import pandas as pd from operator import add d1={ 'Sweden':{'jan':{ '0-5': 5, '6-8': 8, '9-10':19, '11-15': 14, '16-18': 24}, 'march':{ '0-5': 5, '6-8': 18, '9-10': 9, '11-15': 14, '16-18': 24}, 'feb':{ '0-5': 5, '6-8': 7, '9-10': 3, '11-15': 14, '16-18': 24}}, 'Norway':{'jan':{ '0-5': 25, '6-8': 8, '9-10': 45, '11-15': 14, '16-18': 24}, 'march':{ '0-5': 2, '6-8': 8, '9-10': 88, '11-15': 14, '16-18': 24}, 'feb':{ '0-5': 5, '6-8': 48, '9-10': 9, '11-15': 39, '16-18': 24}} }
我目前通过嵌套for循环将其转换为形状为(30,4)的DataFrame,代码如下:
colnames=['country','month','age','revenue'] lst=[] for i in d1.keys(): for j in d1[i].keys(): revenue=list(d1[i][j].items()) l1=list(map(add,[(i,j)]*5,revenue)) lst=lst+l1 df=pd.DataFrame.from_records(lst,columns=colnames)
请问Pandas是否有内置函数可以无需使用循环,以更优雅、高效的方式完成该转换?最快的实现方式是什么?
回答
Pandas提供了多种无需手动嵌套循环的方案,以下是几种常用且高效的实现:
方法1:多层索引展开法(性能最优)
通过构造多层索引结构,再利用Pandas内置的stack()和reset_index()展开,这是速度最快的方式:
# 构造多层索引的DataFrame multi_df = pd.DataFrame.from_dict( {(country, month): ages for country, months in d1.items() for month, ages in months.items()}, orient='index' ) # 展开最内层的age分组并重置索引 df = multi_df.stack().reset_index() # 命名列 df.columns = ['country', 'month', 'age', 'revenue']
这种方法全程使用Pandas的向量化操作,避免了手动循环的内存开销,数据量越大优势越明显。
方法2:生成器结合pd.DataFrame.from_records
用生成器惰性生成数据,代码简洁且内存友好,性能接近多层索引法:
colnames=['country','month','age','revenue'] df = pd.DataFrame.from_records( ( (country, month, age, revenue) for country, months in d1.items() for month, ages in months.items() for age, revenue in ages.items() ), columns=colnames )
生成器不会一次性把所有数据加载到内存,适合处理超大字典。
方法3:pd.json_normalize(可读性优先)
如果更看重代码可读性,json_normalize可以直观展开嵌套结构:
# 先将字典转换为扁平化的列表格式,再用json_normalize处理 flat_data = [ {'country': country, 'month': month, 'age': age, 'revenue': revenue} for country, months in d1.items() for month, ages in months.items() for age, revenue in ages.items() ] df = pd.json_normalize(flat_data)
这种写法逻辑清晰,但性能略逊于前两种方法。
总结
- 最快的实现方式是多层索引展开法,兼顾性能和代码简洁性;
- 生成器法适合内存有限的场景;
json_normalize更适合需要快速理解代码的场景。
内容的提问来源于stack exchange,提问作者Niltzable
相关产品推荐
相关产品推荐

