如何为DataFrame新增MemWeight列存储Weighting值在mem_list中的索引?
解决DataFrame新增列获取列表索引的问题
错误原因
你原来的代码报错是因为mem_list.index()只能接收单个值,而df["Weighting"]是一个Pandas Series(一组数据),直接传进去会让Python无法判断整个Series的布尔值,所以抛出"truth value ambiguous"的错误。
可行解决方案
方法一:字典映射(效率最高,推荐)
先把mem_list转换成元素:索引的映射字典,再用Series.map()批量处理:
# 创建元素到索引的映射字典 weight_to_index = {val: idx for idx, val in enumerate(mem_list)} # 新增MemWeight列 df["MemWeight"] = df["Weighting"].map(weight_to_index)
这种方法处理大数据量时速度最快,因为字典查找是O(1)时间复杂度。
方法二:apply逐元素处理
对Weighting列的每个元素单独调用mem_list.index():
df["MemWeight"] = df["Weighting"].apply(lambda x: mem_list.index(x))
注意:如果Weighting里存在mem_list没有的值,会抛出ValueError,可以加默认值处理:
df["MemWeight"] = df["Weighting"].apply(lambda x: mem_list.index(x) if x in mem_list else -1)
方法三:分类数据类型(适合后续分类操作场景)
利用Pandas的Categorical类型指定顺序,通过codes属性直接获取索引:
df["MemWeight"] = pd.Categorical(df["Weighting"], categories=mem_list, ordered=True).codes
如果元素不在mem_list中,会自动返回-1,无需额外处理。
运行后你会看到,df["MemWeight"]中大部分值为4(对应mem_list里的2),第12行(索引11)的值为1(对应mem_list里的1.97),符合预期。
内容的提问来源于stack exchange,提问作者Robsmith
相关产品推荐
相关产品推荐

