在Pandas中为DataFrame行迭代统计函数:正态分布计算报错
问题描述
我有如下所示的DataFrame,希望新增一行名为Ploss的数据,该行通过以Return行作为均值、Vol行作为标准差的正态分布函数计算得到(即求正态分布在0处的累积分布值)。
尝试调用代码:
Ploss = NormalDist(df.loc['Return'], df.loc['Vol']).cdf(0)
但出现错误,提示sigma必须大于0,而Vol行的数值均为正数。使用单个数值调用时(如Ploss=NormalDist(.11,.14).cdf(0))则可正常运行。
对应的DataFrame:
Iter1 Iter2 Iter3 Iter4 Iter5 Iter6 Iter7 Iter8 Iter9 Iter10 Return 0.04 0.05 0.06 0.07 0.08 0.09 0.10 0.10 0.10 0.11 Vol 0.01 0.02 0.04 0.05 0.07 0.08 0.10 0.11 0.13 0.14
解决方案
核心问题是statistics.NormalDist不支持直接传入pandas Series对象,它仅能处理单个数值。以下两种方法可以解决:
方法一:逐列遍历计算(基于NormalDist)
利用apply方法对DataFrame的每一列单独计算,确保每对均值和标准差都是单个数值:
import pandas as pd from statistics import NormalDist # 假设你的DataFrame变量名为df df.loc['Ploss'] = df.apply(lambda col: NormalDist(col['Return'], col['Vol']).cdf(0), axis=0)
方法二:向量化计算(更高效)
使用scipy.stats.norm的cdf函数,它支持直接传入Series进行向量化运算,无需逐列遍历,计算效率更高:
import pandas as pd from scipy.stats import norm df.loc['Ploss'] = norm.cdf(0, loc=df.loc['Return'], scale=df.loc['Vol'])
内容的提问来源于stack exchange,提问作者wayner
相关产品推荐
相关产品推荐

