如何将Pandas索引元素转为字典值?代码异常排查与解决
问题
需要将DataFrame的索引元素按col1分组生成字典,示例DataFrame如下:
Index_col col1 P1 F1-R1 P2 F1-R1 P3 F1-R1 P4 F1-R1 P5 F1-R2 P6 F1-R2 P7 F1-R2 P8 F1-R2
期望得到的字典:
{'F1-R1': ['P1', 'P2', 'P3', 'P4'], 'F1-R2': ['P5', 'P6', 'P7', 'P8']}
但使用以下代码:
dic = dict.fromkeys(df.col1.unique(), []) for idx, row in df.iterrows(): dic[row["col1"]].append(idx)
得到错误结果:
{'F1-R1': ['P1', 'P2', 'P3', 'P4', 'P5', 'P6', 'P7', 'P8'], 'F1-R2': ['P1', 'P2', 'P3', 'P4', 'P5', 'P6', 'P7', 'P8']}
问题原因
dict.fromkeys(df.col1.unique(), [])会让所有键共享同一个空列表对象。因为Python中列表是可变对象,fromkeys只是把这个列表的引用赋值给每个键,并不是为每个键创建新的空列表。后续调用append时,所有操作都是在同一个列表上进行,最终所有键对应的都是这个被多次修改的列表。
正确实现方法
方法1:使用Pandas内置groupby(推荐)
直接利用Pandas的分组功能,一行代码完成,效率远高于循环:
dic = df.groupby('col1').apply(lambda x: x.index.tolist()).to_dict()
方法2:修复循环逻辑
避免使用fromkeys,为每个键单独创建空列表:
dic = {} for idx, row in df.iterrows(): key = row["col1"] if key not in dic: dic[key] = [] dic[key].append(idx)
方法3:使用collections.defaultdict
借助defaultdict自动为不存在的键创建空列表:
from collections import defaultdict dic = defaultdict(list) for idx, row in df.iterrows(): dic[row["col1"]].append(idx) # 如需转换为普通字典: dic = dict(dic)
内容的提问来源于stack exchange,提问作者Corsair
相关产品推荐
相关产品推荐

