Pandas结合Numpy做多条件列赋值触发ValueError错误如何解决?
报错根因
- Python中位运算符
&的优先级高于比较运算符,原代码中未给条件加括号的写法会被优先计算3 & df['Value'].astype(int),逻辑完全偏离预期,最终触发Series布尔值歧义的报错。
修复方法
给两个比较条件分别用括号包裹,明确计算优先级即可,修复后完整代码如下:
# Import pandas library import pandas as pd import numpy as np # initialize list of lists data = [[1, 0], [4, 0], [8, 0]] # Create the pandas DataFrame df = pd.DataFrame(data, columns=['Value', 'Test']) df['Test'] = np.where((df['Value'].astype(int) >= 3) & (df['Value'].astype(int) <= 7), 1, 2) # print dataframe. print(df)
简化写法
也可以直接使用pandas自带的between方法实现闭区间判断,代码更简洁易读:
df['Test'] = np.where(df['Value'].astype(int).between(3, 7), 1, 2)
运行输出
Value Test 0 1 2 1 4 1 2 8 2
内容的提问来源于stack exchange,提问作者JerryC
相关产品推荐
相关产品推荐

