Pandas 2.1.4重命名DataFrame时遇TypeError: unhashable type: 'set'
问题
我有一个大型DataFrame(news1),想要根据另一个DataFrame(dic)中的值,在列名匹配的前提下重命名news1的列。我编写了如下嵌套循环代码:
for index,col in enumerate(news1.columns): for index_dic,col_dic in enumerate(dic.columns): if col==col_dic: print(col,dic.iloc[1, index_dic]) print(type(col)) print(type(dic.iloc[1, index_dic])) news1.rename(columns={f'{col}':f'{dic.iloc[1, index_dic]}'},inplace=True)
运行时抛出TypeError: unhashable type: 'set'错误,报错回溯信息如下:
----> 1 news1.rename(columns={'c15.10':'WA_ambiguous-expectation'}) File c:\Users\\anaconda3\envs\PythonCourse2023\Lib\site-packages\pandas\core\frame.py:5518, in DataFrame.rename(self, mapper, index, columns, axis, copy, inplace, level, errors) 5399 def rename( 5400 self, 5401 mapper: Renamer | None = None, (...) 5409 errors: IgnoreRaise = "ignore", 5410 ) -> DataFrame | None: 5411 """ 5412 Rename columns or index labels. 5413 (...) 5516 4 3 6 5517 """ -> 5518 return super()._rename( 5519 mapper=mapper, 5520 index=index, 5521 columns=columns, 5522 axis=axis, 5523 copy=copy, 5524 inplace=inplace, 5525 level=level, 5526 errors=errors, 5527 ) File c:\Users\\anaconda3\envs\PythonCourse2023\Lib\site-packages\pandas\core\generic.py:1086, in NDFrame._rename(self, mapper, index, columns, axis, copy, inplace, level, errors) 1079 missing_labels = [ 1080 label 1081 for index, label in enumerate(replacements) 1082 if indexer[index] == -1 1083 ] 1084 raise KeyError(f"{missing_labels} not found in axis") -> 1086 new_index = ax._transform_index(f, level=level) 1087 result._set_axis_nocheck(new_index, axis=axis_no, inplace=True, copy=False) 1088 result._clear_item_cache() File c:\Users\\anaconda3\envs\PythonCourse2023\Lib\site-packages\pandas\core\indexes\base.py:6465, in Index._transform_index(self, func, level) 6463 return type(self).from_arrays(values) 6464 else: -> 6465 items = [func(x) for x in self] 6466 return Index(items, name=self.name, tupleize_cols=False) File c:\Users\\anaconda3\envs\PythonCourse2023\Lib\site-packages\pandas\core\indexes\base.py:6465, in <listcomp>(.0) 6463 return type(self).from_arrays(values) 6464 else: -> 6465 items = [func(x) for x in self] 6466 return Index(items, name=self.name, tupleize_cols=False) File c:\Users\anaconda3\envs\PythonCourse2023\Lib\site-packages\pandas\core\common.py:507, in get_rename_function.<locals>.f(x) 506 def f(x): --> 507 if x in mapper: 508 return mapper[x] 509 else:
打印结果显示列名和目标重命名值均为str类型:
c15.10 WA_ambiguous-expectation <class 'str'> <class 'str'>
无法定位报错原因,寻求帮助。
解决方案
错误根源
报错提示unhashable type: 'set',说明传入rename的字典中存在集合类型的键或值。虽然你打印的单个案例是字符串,但dic中其他列/行的值可能是集合;另外,嵌套循环中多次调用inplace=True的rename,会实时修改news1的列名,导致后续循环匹配出现异常,同时这种方式对大型DataFrame效率极低。
修复代码
步骤1:构建完整映射字典
先一次性生成所有需要重命名的列名映射,避免循环中频繁修改DataFrame:
# 提取dic第1行作为新列名(注意iloc[1]对应第二行,若需第一行请改用iloc[0]) new_name_series = dic.iloc[1] # 构建仅包含news1现有列的映射字典 rename_map = {col: new_name_series[col] for col in news1.columns if col in new_name_series.index}
步骤2:一次性完成重命名
用构建好的字典执行一次重命名操作:
news1.rename(columns=rename_map, inplace=True)
简化优化
如果dic的列名和news1的列名完全对应,可直接替换列名:
# 仅保留news1列对应的新名称 valid_new_names = dic.iloc[1][news1.columns] news1.columns = valid_new_names.values
内容的提问来源于stack exchange,提问作者Mostafa Bouzari
相关产品推荐
相关产品推荐

