Pandas:如何用列表推导式或更优雅方法基于两列条件创建新列
问题描述
给出的DataFrame定义如下:
import pandas as pd df = pd.DataFrame ({ 'probability': [0.51, 0.48, 0.52, 0.71, 0.38, 0.22, 0.59, 0.70, 0.44, 0.62, 0.38], 'indicator': [1, 0, 0, 1, 0, 1, 0, 1, 0, 1, 0] })
需求:新增accuracy列,当probability>0.5且indicator=1时取值1,否则为0。已用np.select实现但觉得繁琐,尝试其他写法时出现报错:
TypeError: Cannot perform 'rand_' with a dtyped [int64] array and scalar of type [bool]
询问是否有更优雅的实现方式(如列表推导式)。
解决方案
1. Pandas向量化布尔运算(推荐,效率最高)
利用Pandas原生的向量化操作,直接组合布尔条件后转换为整数(True对应1,False对应0):
df['accuracy'] = ((df['probability'] > 0.5) & (df['indicator'] == 1)).astype(int)
也可以用np.where实现相同逻辑:
import numpy as np df['accuracy'] = np.where((df['probability'] > 0.5) & (df['indicator'] == 1), 1, 0)
2. 列表推导式实现
如果偏好迭代式写法,可通过zip配对两列数据,逐个判断生成结果:
df['accuracy'] = [1 if p > 0.5 and i == 1 else 0 for p, i in zip(df['probability'], df['indicator'])]
报错原因说明
你遇到的TypeError是因为两个问题:一是误用了Python原生的and运算符(仅支持单个布尔值判断),而非Pandas/NumPy的向量运算符&;二是未给每个条件单独加括号,因运算符优先级问题导致0.5 & df['indicator']这类错误运算(int类型标量与bool数组无法执行按位与操作)。只需用&替代and,并为每个条件添加括号即可避免报错。
内容的提问来源于stack exchange,提问作者equanimity
相关产品推荐
相关产品推荐

