DataFrame真值歧义报错:基于sales列生成新列失败求助
解决Pandas中"The truth value of a DataFrame is ambiguous"报错问题
报错信息
The truth value of a DataFrame is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all()
数据集
| abd | sales |
|---|---|
| atf | 1.2 |
| adstys | 0.9 |
| tyugug | 5.6 |
| gfygv | 1.1 |
需求
基于sales列创建新列new_sales_index:
- 当
sales值大于1.2时,设为1 - 当
sales值小于1.2时,设为0
问题代码
df['new_sales_index']=df['sales'].apply(lambda x: [1 if y > 1.2 else 0])
错误原因
- lambda函数误用了未定义的变量
y,apply传入的x已经是sales列的单个float元素,不需要额外变量,逻辑混乱触发真值判断歧义错误。 - 函数返回列表类型,但需求只需要单个数值,类型不匹配加剧了问题。
解决方案
方法1:修正apply的lambda逻辑
直接用传入的x做判断,返回单个数值:
df['new_sales_index'] = df['sales'].apply(lambda x: 1 if x > 1.2 else 0)
方法2:Pandas矢量化操作(推荐,效率更高)
避免使用apply,直接对整列做矢量化判断,再转换为整数:
df['new_sales_index'] = (df['sales'] > 1.2).astype(int)
说明:df['sales'] > 1.2生成布尔值Series,astype(int)自动将True转为1,False转为0,完全符合需求。
方法3:使用numpy.where
借助numpy条件判断函数实现:
import numpy as np df['new_sales_index'] = np.where(df['sales'] > 1.2, 1, 0)
处理后结果
| abd | sales | new_sales_index |
|---|---|---|
| atf | 1.2 | 0 |
| adstys | 0.9 | 0 |
| tyugug | 5.6 | 1 |
| gfygv | 1.1 | 0 |
内容的提问来源于stack exchange,提问作者nuke
相关产品推荐
相关产品推荐

