Pandas根据动态列存在性更新列触发TypeError问题求助
问题分析与解决方案
错误原因
你代码的核心问题是:"independant_value_" + the_dataframe["suffix"].str.lower()生成的是Series对象,而in the_dataframe.columns是用来判断单个字符串是否在列名集合中的操作,无法直接对整个Series执行该判断——因为Series是不可哈希的类型,所以触发TypeError: unhashable type: 'Series'。
另外注意:你代码里同时出现了independant_value_和independent_value_(拼写差异,少了一个字母e),这可能导致后续列名匹配失败,建议先统一拼写。
解决方案
方案一:逐行处理(逻辑直观,小数据量适用)
通过apply逐行检查目标列是否存在,再计算对应值:
# 统一列名前缀,修正拼写问题 prefix = "independent_value_" # 生成每行对应的目标列名 target_cols = prefix + the_dataframe["suffix"].str.lower() # 逐行判断并赋值 the_dataframe["dependant_value"] = the_dataframe.apply( lambda row: row["another_column"] * row[target_cols.loc[row.name]] if target_cols.loc[row.name] in the_dataframe.columns else 0, axis=1 )
方案二:批量处理(效率更高,大数据量推荐)
先筛选出目标列存在的行,再批量计算赋值,避免逐行循环的性能损耗:
prefix = "independent_value_" target_cols = prefix + the_dataframe["suffix"].str.lower() # 标记哪些行的目标列存在于DataFrame中 cols_exist = target_cols.isin(the_dataframe.columns) # 先初始化结果列为0 the_dataframe["dependant_value"] = 0 # 对目标列存在的行,批量计算乘积 mask = cols_exist the_dataframe.loc[mask, "dependant_value"] = ( the_dataframe.loc[mask, "another_column"] * the_dataframe.loc[mask, target_cols[mask].values] )
内容的提问来源于stack exchange,提问作者HuLu ViCa
相关产品推荐
相关产品推荐

