Pandas应用自定义函数生成match_status列报错,求正确实现方式
Pandas新增match_status列的正确实现方式
问题背景
现有如下Pandas DataFrame:
home_team_goal away_team_goal id 1 1 1 2 0 0 3 0 3 4 5 0 5 1 3
需要新增match_status列,逻辑由以下函数定义:
def match_status(home_goals, other_goals): if home_goals > other_goals: return 'WIN' elif home_goals < other_goals: return 'LOSE' else: return 'DRAW'
使用语句df_match.apply(match_status, 1, df_match['home_team_goal'], df_match['away_team_goal'])时,触发如下ValueError:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-83-2e2115165ea9> in <module>() ----> 1 df_match.apply(match_status, 1, df_match['home_team_goal'], df_match['away_team_goal']) 3 frames /usr/local/lib/python3.7/dist-packages/pandas/core/generic.py in __nonzero__(self) 1536 def __nonzero__(self): 1537 raise ValueError( -> 1538 f"The truth value of a {type(self).__name__} is ambiguous. " 1539 "Use a.empty, a.bool(), a.item(), a.any() or a.all()." 1540 ) ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
错误原因
原调用方式直接将整个Series传入match_status函数,函数内的比较操作会返回布尔类型的Series,而if语句无法判断整个布尔Series的真假,因此抛出歧义错误。
正确实现方法
方法一:修改函数接收行对象
调整函数,让它接收每行数据并提取对应列的值:
def match_status(row): home_goals = row['home_team_goal'] other_goals = row['away_team_goal'] if home_goals > other_goals: return 'WIN' elif home_goals < other_goals: return 'LOSE' else: return 'DRAW' # 应用函数并新增列 df_match['match_status'] = df_match.apply(match_status, axis=1)
方法二:保留原函数,用Lambda包装传递参数
通过Lambda函数将每行的对应列值传入原函数:
df_match['match_status'] = df_match.apply( lambda row: match_status(row['home_team_goal'], row['away_team_goal']), axis=1 )
方法三:矢量化操作(效率更高)
避免使用apply,直接用Pandas/Numpy的矢量化函数实现,适合大数据量场景:
import numpy as np # 嵌套np.where实现 df_match['match_status'] = np.where( df_match['home_team_goal'] > df_match['away_team_goal'], 'WIN', np.where( df_match['home_team_goal'] < df_match['away_team_goal'], 'LOSE', 'DRAW' ) ) # 或者用np.select实现(逻辑更清晰) conditions = [ df_match['home_team_goal'] > df_match['away_team_goal'], df_match['home_team_goal'] < df_match['away_team_goal'], df_match['home_team_goal'] == df_match['away_team_goal'] ] choices = ['WIN', 'LOSE', 'DRAW'] df_match['match_status'] = np.select(conditions, choices)
内容的提问来源于stack exchange,提问作者Indika Rajapaksha
相关产品推荐
相关产品推荐

