Pandas 0.20与0.24用lambda生成字典的行为差异及适配需求
Pandas 0.20.1与0.24.2 apply返回结果差异问题及解决方法
问题背景
使用Python 2.7.13,发现Pandas 0.20.1与0.24.2版本在apply结合lambda函数生成字典时存在行为差异。执行以下代码:
import pandas as pd df = pd.DataFrame({'cluster': ['5', '5', '5', '5', '5', '5'], 'mdse_item_i': ['23627102', '23627102', '23627102', '23627102', '23627102', '23627102'], 'predPriceQty': ['35.675543', '35.675543', '35.675543', '35.675543', '35.675543', '35.675543'], 'schedule_i': ['56', '56', '56', '56', '56', '56'], 'segment_id': ['4123', '4123', '4144', '4161', '4295', '4454'], 'wk': ['1', '1', '1', '1', '1', '1']} ) df.set_index(['segment_id', 'cluster'], inplace=True) df.apply(lambda row: {row['schedule_i']: {row['mdse_item_i']: {row['wk']: row['predPriceQty']}}}, axis=1)
- Pandas 0.24.2返回包含嵌套字典的Series:
segment_id cluster 4123 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} 4144 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} 4161 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} 4295 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} 4454 5 {u'56': {u'23627102': {u'1': u'35.675543'}}} dtype: object
- Pandas 0.20.1返回全NaN的DataFrame:
mdse_item_i predPriceQty schedule_i wk segment_id cluster 4123 5 NaN NaN NaN NaN 5 NaN NaN NaN NaN 4144 5 NaN NaN NaN NaN 4161 5 NaN NaN NaN NaN 4295 5 NaN NaN NaN NaN 4454 5 NaN NaN NaN NaN
最终目标是生成以索引为键、嵌套字典为值的目标字典,且需处理10万级数据,无法升级Pandas版本,需在0.20.1中实现。
差异原因
Pandas 0.20.1中,当apply(axis=1)返回字典时,默认会将字典的键当作新DataFrame的列名,尝试将值对应到列中。但此处lambda返回的字典键是schedule_i的取值(如'56'),与原DataFrame的列名不匹配,因此所有列都填充为NaN。
而Pandas 0.24.2及之后版本调整了该行为:当apply返回字典时,会将整个字典作为单个元素存入Series,不再自动扩展为DataFrame。
0.20.1版本解决方案
以下提供三种高效实现方式,均避免低效的行迭代,适配10万级数据:
方法1:使用result_type='reduce'参数
通过指定result_type='reduce',强制apply返回Series而非自动扩展为DataFrame:
result_series = df.apply(lambda row: {row['schedule_i']: {row['mdse_item_i']: {row['wk']: row['predPriceQty']}}}, axis=1, result_type='reduce') target_dict = result_series.to_dict()
方法2:返回单元素Series后提取值
让lambda返回仅包含目标字典的单元素Series,再提取该列得到结果Series:
result_series = df.apply(lambda row: pd.Series([{row['schedule_i']: {row['mdse_item_i']: {row['wk']: row['predPriceQty']}}]}), axis=1)[0] target_dict = result_series.to_dict()
方法3:使用itertuples+map(性能最优)
itertuples比apply的行处理更快,结合map批量生成字典,最后用zip关联索引:
target_dict = dict(zip(df.index, map(lambda t: {t.schedule_i: {t.mdse_item_i: {t.wk: t.predPriceQty}}}, df.itertuples(index=False))))
内容的提问来源于stack exchange,提问作者Vivek
相关产品推荐
相关产品推荐

