如何在Pandas中基于下一行值创建新列且避免布尔逻辑报错
代码修复方案
错误原因
你代码的核心错误是直接对整个Pandas Series对象使用布尔判断,Python无法判断一整列布尔值的整体真值,因此触发ValueError: The truth value of a Series is ambiguous报错。同时你使用列表推导式的写法没有逐行遍历数据,逻辑也不成立。
修复后完整代码
import pandas as pd import numpy as np # 原有逻辑保留 full_count['batter-pitcher'] = full_count['batter'].astype(str) + "---" + full_count['pitcher'].astype(str) full_count = full_count.loc[(full_count['balls'] == 3) & (full_count['strikes'] == 2)] # 新增列逻辑,无需额外创建prior、prior1中间列 full_count['prev_pitch'] = np.where( full_count['batter-pitcher'] == full_count['batter-pitcher'].shift(-1), full_count['description'].shift(-1), 'na' )
代码说明
full_count['batter-pitcher'].shift(-1):获取当前行下一行的batter-pitcher值np.where为向量化判断函数,会逐行校验匹配条件:- 条件为真(当前行和下一行
batter-pitcher值匹配):返回下一行的description字段值 - 条件为假:返回
na
- 条件为真(当前行和下一行
- 整体为Pandas向量化操作,执行效率远高于逐行遍历的写法
如果你不想引入numpy依赖,也可以用Pandas原生where方法实现:
full_count['prev_pitch'] = full_count['description'].shift(-1).where( full_count['batter-pitcher'] == full_count['batter-pitcher'].shift(-1), other='na' )
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

