如何将自定义Python函数应用到pandas DataFrame列并解决ValueError报错
问题说明
我编写了如下函数,可接收四个输入值并基于输入计算返回结果:
def python_function(a, b, c, d): if [a, b, c, d].count(0) == 4: return "NA" average = (a + b + c + d) / (4 - [a, b, c, d].count(0)) # 计算q1改a,q2改b,q3改c,q4改d if c >= average: if c > b: return "G" else: return "S" elif c < average: return "B" return "NA"
函数调用示例:
python_function(5.3,9.7,.4,0) # 返回结果:'B' python_function(5.3,9.7,10.4,0) # 返回结果:'G'
将该函数直接应用于pandas DataFrame的列时触发了报错,已知需要适配逻辑运算符的浮点值实现方法,不清楚具体操作步骤。
所用DataFrame示例:
q1_profit q2_profit q3_profit q4_profit 0 89969.7 112896.7 25665.4 0 1 1.6 459.9 295.9 0 2 0.9 9.5 5.3 0 3 1396.1 1105.2 0.2 0 4 17.9 365.5 191.1 0
字段数据类型:
q1_profit 1600 non-null float64 q2_profit 1600 non-null float64 q3_profit 1600 non-null float64 q4_profit 1600 non-null int64
调用写法:
data["rating"] = python_function(data["q1_profit"],data["q2_profit"],data["q3_profit"],data["q4_profit"])
报错信息:
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-51-6dba2870dd9c> in <module> ----> 1 data["rating"] = python_function(data["q1_profit"],data["q2_profit"],data["q3_profit"],data["q4_profit"]) <ipython-input-39-47792387b172> in python_function(a, b, c, d) 1 def python_function(a, b, c, d): ----> 2 if [a, b, c, d].count(0) == 4: 3 return "NA" 4 5 average = (a + b + c + d) / (4 - [a, b, c, d].count(0)) ~\Anaconda3\lib\site-packages\pandas\core\generic.py in __nonzero__(self) 1476 raise ValueError("The truth value of a {0} is ambiguous. " 1477 "Use a.empty, a.bool(), a.item(), a.any() or a.all()." -> 1478 .format(self.__class__.__name__)) 1479 1480 __bool__ = __nonzero__ ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
报错原因
现有函数是针对单个标量值设计的,直接传入整列Series对象时,if判断需要单个布尔值,但返回的是一整列布尔结果,pandas无法判定整列的真值,因此抛出该错误。
解决方案
方案1:逐行调用原函数
不需要修改原有函数逻辑,通过apply按行遍历传参即可,适合万行以下的小数据集:
data["rating"] = data.apply(lambda row: python_function(row["q1_profit"], row["q2_profit"], row["q3_profit"], row["q4_profit"]), axis=1)
方案2:向量化实现(性能更高)
重写为pandas原生向量化操作,避免逐行遍历的性能损耗,适合十万行以上的大数据集:
import numpy as np # 计算每行0的个数 zero_cnt = (data[["q1_profit", "q2_profit", "q3_profit", "q4_profit"]] == 0).sum(axis=1) # 计算非零平均值 data["average"] = data[["q1_profit", "q2_profit", "q3_profit", "q4_profit"]].sum(axis=1) / (4 - zero_cnt) # 多条件匹配赋值 conditions = [ zero_cnt == 4, data["q3_profit"] >= data["average"], data["q3_profit"] < data["average"] ] choices = [ "NA", np.where(data["q3_profit"] > data["q2_profit"], "G", "S"), "B" ] data["rating"] = np.select(conditions, choices, default="NA") # 可选:删除中间生成的辅助列 data.drop("average", axis=1, inplace=True)
内容的提问来源于stack exchange,提问作者sonu_db_work
相关产品推荐
相关产品推荐

