使用np.select基于pandas两列新增列结果不符合预期如何解决
问题原因
- 运算符优先级错误:Python中按位与
&的优先级高于比较运算符==,你编写的data['B']==1 & data['current_pair'].str.contains('Emo/', na=False)实际执行顺序为data['B'] == (1 & data['current_pair'].str.contains('Emo/', na=False)),逻辑完全不符合预期,导致条件判断失效。 - 变量误用:你定义的DataFrame变量名为
df,但编写条件时用的是原始字典变量data,字典没有DataFrame列的str.contains等方法,也会导致逻辑错误。
解决方法
给每个比较条件单独加括号明确运算顺序,同时修正条件引用的变量为正确的DataFrame名即可,修正后的完整代码如下:
import pandas as pd import numpy as np data = {'current_pair': ['"["StimusNeu/2357.jpg","StimusNeu/5731.jpg"]"', '"["StimusEmo/6350.jpg","StimusEmo/3230.jpg"]"', '"["StimusEmo/3215.jpg","StimusEmo/9570.jpg"]"','"["StimusNeu/7020.jpg","StimusNeu/7547.jpg"]"', '"["StimusNeu/7080.jpg","StimusNeu/7179.jpg"]"'], 'B': [1, 0, 1, 1, 0] } df = pd.DataFrame(data) # 修正后条件,每个比较逻辑单独用括号包裹 conditions=[ (df['B']==1) & (df['current_pair'].str.contains('Emo/', na=False)), (df['B']==1) & (df['current_pair'].str.contains('Neu/', na=False)), df['B']==0 ] choices = [0, 1, 2] df['C'] = np.select(conditions, choices, default=np.nan)
运行上述代码后得到的结果和你预期的输出完全一致。
内容的提问来源于stack exchange,提问作者Catherine
相关产品推荐
相关产品推荐

