如何用Pandas apply为DataFrame每行计算Z-spread?报错解决
Z-Spread计算:apply方法替代for循环的问题解决
问题背景
我有一个包含数千行的DataFrame,已定义zspread_solution和fun_solve两个函数,尝试通过df.apply(fun_solve, axis=1)为每行计算Z-spread时触发TypeError:fsolve的func参数输入输出形状不匹配(应为(6,)实际为(1050,))。目前已通过for循环+loc成功实现计算,但希望了解如何用apply完成该需求。
错误代码
def zspread_solution(x): FV = x[1] COUPON_RATE = x[2] T = x[3] N = x[4] C = COUPON_RATE * T * N fun = - FV for item_index in range(int(x[5] + 1)): if item_index != x[5]: DELTA_T = (item_index + 1) * T rate = df[item_index + 1] + x[0] NUMERATOR = (1 + rate)**DELTA_T item = C/NUMERATOR fun += item else: DELTA_T = item_index * T rate = df[item_index] + x[0] NUMERATOR = (1 + rate)**DELTA_T item = N/NUMERATOR fun += item return fun def fun_solve(df): closing_price = df['closing_price'] coupon_rate = df['coupon_rate'] Interest_payment_interval = df['Interest_payment_interval'] face_value = df['face_value'] Interest_payment_time = df['Interest_payment_time'] x = [0, closing_price, coupon_rate, Interest_payment_interval, face_value, Interest_payment_time] return fsolve(zspread_solution, x)[0]
报错信息
TypeError: fsolve: there is a mismatch between the input and output shape of the 'func' argument 'zspread_solution'.Shape should be (6,) but it is (1050,)
已实现的for循环代码
def fu(x): fun = - FV for item_index in range(int(TIM + 1)): if item_index != TIM: DELTA_T = (item_index + 1) * T rate = df.loc[0, item_index + 1] + x NUMERATOR = (1 + rate)**DELTA_T item = C/NUMERATOR fun += item else: DELTA_T = item_index * T rate = df.loc[0, item_index] + x NUMERATOR = (1 + rate)**DELTA_T item = N/NUMERATOR fun += item return fun zsp = [] for index in df.index: FV = df.loc[index, 'closing_price'] COUPON_RATE = df.loc[index, 'coupon_rate'] T = df.loc[index, 'Interest_payment_interval'] N = df.loc[index, 'face_value'] C = COUPON_RATE * T * N TIM = df.loc[index, 'Interest_payment_time'] item = fsolve(fu, 0)[0] zsp.append(item) df['zspread'] = zsp
测试数据
closing_price:107.7301, 106.0029, 105.1495, 105.2768 coupon_rate:4.39, 3.65, 3.6, 3.66 Interest_payment_time:8, 20, 20, 15 Interest_payment_interval:1, 1, 1, 1 face_value:100, 100, 100, 100 0:0.01342 1:0.021977 2:0.022843 3:0.023907 4:0.02465 5:0.025296 6:0.026264 7:0.0268 8:0.026793 9:0.026781 10:0.026776 11:0.026875 12:0.027104 13:0.0274 14:0.027702 15:0.027947 16:0.028099 17:0.028181 18:0.028221 19:0.028247 20:0.028288 21:0.028368 22:0.028487
问题分析与修正方案
错误原因
- return语句位置错误:
zspread_solution的return写在for循环内部,导致第一次循环就直接返回,未完成所有现金流折现计算。 - fsolve优化变量冗余:
fun_solve传给fsolve的初始值是长度为6的列表,但实际仅需优化Z-spread一个变量,其余参数是固定的行数据,不需要作为优化变量。 - 全局变量依赖:直接引用全局
df导致计算时取到整列数据,返回Series而非单个数值,引发形状不匹配。
修正后的apply实现代码
from scipy.optimize import fsolve import pandas as pd # 仅将Z-spread作为优化变量,其余参数通过args传入 def zspread_solution(z_spread, FV, COUPON_RATE, T, N, TIM, rate_df): C = COUPON_RATE * T * N fun = -FV for item_index in range(int(TIM + 1)): if item_index != TIM: delta_t = (item_index + 1) * T rate = rate_df[item_index + 1] + z_spread numerator = (1 + rate) ** delta_t fun += C / numerator else: delta_t = item_index * T rate = rate_df[item_index] + z_spread numerator = (1 + rate) ** delta_t fun += N / numerator # 把return移到循环外部 return fun # 处理单行数据的函数,通过参数传递利率数据 def fun_solve(row, rate_df): closing_price = row['closing_price'] coupon_rate = row['coupon_rate'] interval = row['Interest_payment_interval'] face_value = row['face_value'] payment_time = row['Interest_payment_time'] # 仅优化Z-spread一个变量,初始值设为0 z_spread = fsolve(zspread_solution, x0=0, args=(closing_price, coupon_rate, interval, face_value, payment_time, rate_df))[0] return z_spread # 构造测试DataFrame data = { 'closing_price': [107.7301, 106.0029, 105.1495, 105.2768], 'coupon_rate': [4.39, 3.65, 3.6, 3.66], 'Interest_payment_time': [8, 20, 20, 15], 'Interest_payment_interval': [1, 1, 1, 1], 'face_value': [100, 100, 100, 100], 0: [0.01342]*4, 1: [0.021977]*4, 2: [0.022843]*4, 3: [0.023907]*4, 4: [0.02465]*4, 5: [0.025296]*4, 6: [0.026264]*4, 7: [0.0268]*4, 8: [0.026793]*4, 9: [0.026781]*4, 10: [0.026776]*4, 11: [0.026875]*4, 12: [0.027104]*4, 13: [0.0274]*4, 14: [0.027702]*4, 15: [0.027947]*4, 16: [0.028099]*4, 17: [0.028181]*4, 18: [0.028221]*4, 19: [0.028247]*4, 20: [0.028288]*4, 21: [0.028368]*4, 22: [0.028487]*4 } df = pd.DataFrame(data) # 使用apply计算,传入利率数据列 df['zspread'] = df.apply(fun_solve, axis=1, rate_df=df)
关键修正点
- 优化变量单一化:fsolve仅针对Z-spread进行优化,其余参数通过
args传递,避免形状不匹配。 - 修复return位置:确保
zspread_solution完成所有循环计算后再返回结果。 - 消除全局依赖:通过参数传递所需数据,避免引用全局变量导致的Series返回问题。
内容的提问来源于stack exchange,提问作者1670511081
相关产品推荐
相关产品推荐

