Python Pandas自定义多参数函数.apply()调用问题及报错求助
自定义函数应用到DataFrame的问题解决
问题背景
需要将自定义函数应用到DataFrame,函数接收DataFrame的两列作为参数:
函数定义
def faustmann(profit, age): interest = 1.03 cost = 262 revenue = (profit - ((cost * interest) ** age)) / ((interest ** age) - 1) return revenue
目标DataFrame
Year age Total_MCuFt Loblolly_MCuFt Total_SCuFt Loblolly_SCuFt pulp revenue 0 2026 0 0.0 0.0 0.0 0.0 0.0 0.000 1 2031 5 0.0 0.0 0.0 0.0 0.0 0.000 2 2036 10 214.1 214.1 0.0 0.0 214.1 57.807 3 2041 15 1908.8 1908.8 0.0 0.0 1908.8 515.376 4 2046 20 3418.4 3418.4 0.0 0.0 3418.4 922.968
尝试过的方法及错误
使用
args参数pine_800["faustmann"] = pine_800.apply(faustmann, args=(pine_800.revenue, pine_800.age), axis=1)报错:
TypeError: faustmann() takes 2 positional arguments but 3 were given直接传入列作为参数
pine_800["faustmann"] = pine_800.apply(faustmann(profit= pine_800.revenue, age=pine_800.age), axis=1)报错:
AssertionError使用lambda函数
pine_800["faustmann"] = pine_800.apply(lambda x: faustmann(x["revenue"], x["age"]), axis=1)触发警告:
RuntimeWarning: overflow encountered in double_scalars,返回无限/错误值直接在lambda中写计算逻辑
pine_800["faustmann"] = pine_800.apply(lambda x: (pine_800.revenue -(cost*interest**pine_800.age)/(interest*pine_800.age)-1), axis=1)报错:
ValueError: Expected a 1D array, got an array with shape (40, 40)使用
np.vectorizepine_800["faustmann"] = np.vectorize(faustmann)(pine_800.revenue, age=pine_800.age)未解决除零错误,仍触发
ZeroDivisionError
额外问题:部分行profit为0,所有方法均触发ZeroDivisionError,np.seterr(divide='ignore')无效,需忽略错误或填充NaN/"div/0"。
解决方案
1. 修复函数,处理异常逻辑
先修改函数,加入对除零、无效输入的处理,避免报错:
import numpy as np def faustmann(profit, age): interest = 1.03 cost = 262 # 处理分母为0的情况(age=0时,interest**age=1,分母为0) if age == 0 or np.isclose((interest ** age) - 1, 0): return np.nan # 处理profit为0的情况,可根据需求调整返回值 if profit == 0: return np.nan try: revenue = (profit - ((cost * interest) ** age)) / ((interest ** age) - 1) return revenue except (ZeroDivisionError, OverflowError): return np.nan
2. 正确应用函数到DataFrame
方法A:apply+lambda(简单直观)
修正后的lambda调用能处理异常,返回合法值或NaN:
pine_800["faustmann"] = pine_800.apply(lambda x: faustmann(x["revenue"], x["age"]), axis=1)
方法B:向量化运算(高效推荐)
对于大数据集,向量化运算效率远高于逐行apply:
import numpy as np interest = 1.03 cost = 262 # 计算分母,避免除零 denominator = (interest ** pine_800["age"]) - 1 # 标记需要设为NaN的异常行(age=0、分母为0、profit为0) mask = (pine_800["age"] == 0) | np.isclose(denominator, 0) | (pine_800["revenue"] == 0) # 批量计算结果 pine_800["faustmann"] = (pine_800["revenue"] - ((cost * interest) ** pine_800["age"])) / denominator # 异常位置填充NaN pine_800.loc[mask, "faustmann"] = np.nan
方法C:np.vectorize(兼容原函数调用)
使用处理过异常的函数配合np.vectorize:
pine_800["faustmann"] = np.vectorize(faustmann)(pine_800["revenue"], pine_800["age"])
3. 替换NaN值(可选)
如果需要将NaN替换为指定内容,比如"div/0":
pine_800["faustmann"] = pine_800["faustmann"].fillna("div/0")
注意:替换后列类型会变为object,若需保持数值类型,建议保留NaN。
内容的提问来源于stack exchange,提问作者krawall
相关产品推荐
相关产品推荐

