如何在pandas和numpy中基于多条件生成新列并解决报错问题
pandas多条件判断赋值报错解决方案及高效实现
报错原因
你触发错误的核心问题是df1['country'] in allowed_countries的写法不符合pandas语法规则:in关键字无法直接对pandas Series对象做逐元素的成员判断,会尝试将整个序列转换为单个布尔值,因此抛出The truth value of a Series is ambiguous错误。你需要用pandas内置的.isin()方法实现逐元素的成员校验。
修正后可运行代码
import pandas as pd import numpy as np # 原始构造数据 df1 = pd.DataFrame({'name':['Sara', 'John', 'Christine'], 'country': ['US', 'UK', 'CA'], 'Age': [10,20,40]}) allowed_countries = ['UK','India','Germany'] # 修正条件:用.isin()做成员匹配,同时匹配需求中的Age>10规则 conditions = [ np.logical_and(df1['Age'] > 10 , df1['country'].isin(allowed_countries)), np.logical_and(df1['Age'] > 10, df1['country'] == 'CA') ] choices = [ "Allowed", 'Partially allowed'] df1['admission'] = np.select(conditions, choices, default=np.nan)
运行结果
最终得到的df1如下:
name country Age admission 0 Sara US 10 NaN 1 John UK 20 Allowed 2 Christine CA 40 Partially allowed
如果你的实际规则是Age>=10,只需将条件中的df1['Age'] > 10替换为df1['Age'] >= 10即可。
性能说明
该方案是完全向量化的操作,无逐行遍历逻辑,5万行数据执行耗时在毫秒级,远优于apply+lambda的逐行处理方案,完全满足性能需求。
内容的提问来源于stack exchange,提问作者MTALY
相关产品推荐
相关产品推荐

