Scipy lognorm拟合无法收敛至正确sigma参数问题求助
解决scipy拟合对数正态分布的收敛异常问题
问题背景
手动设定sigma=0.15、mu=2时,能正常画出对数正态分布的拟合曲线,但使用curve_fit拟合直方图数据时,sigma收敛至极小值(~7e-5),协方差矩阵全为inf,即使给定初始猜测[2, 0.15]也无法得到正确结果。
原拟合代码及结果:
from scipy.optimize import curve_fit y, x = np.histogram(z_data, bins=100) x = (x[1:]+x[:-1])/2 # bin centers def fit_lognormal(x, mu, sigma): return lognorm.pdf(x, sigma, scale=np.exp(mu)) p0 = [2, 0.15] params = curve_fit(fit_lognormal, x, y, p0=p0) params
返回结果:
(array([1.97713319e+00, 6.98238900e-05]), array([[inf, inf], [inf, inf]]))
问题原因
- 量级不匹配:直方图的
y是频数,而lognorm.pdf输出的是概率密度,两者量级差异极大,导致最小二乘拟合时权重失衡。 - 数值不稳定:当
sigma趋近于0时,对数正态PDF会变得异常尖锐,拟合过程中容易出现数值溢出或梯度消失,引发协方差矩阵为inf。 - 方法选择不当:
curve_fit的普通最小二乘并非分布拟合的最优方法,极大似然估计(MLE)更适合这类问题。
解决方案
方案1:将直方图转换为频率密度后拟合
把直方图的频数转换为频率密度(频数 / (组距 × 样本总数)),让其与PDF的量级匹配,同时为sigma设置下界避免趋近于0:
from scipy.stats import lognorm from scipy.optimize import curve_fit import numpy as np y, x_bins = np.histogram(z_data, bins=100) bin_width = x_bins[1] - x_bins[0] x = (x_bins[1:] + x_bins[:-1])/2 # 转换为频率密度 y_density = y / (bin_width * len(z_data)) def fit_lognormal(x, mu, sigma): return lognorm.pdf(x, sigma, scale=np.exp(mu)) p0 = [2, 0.15] # 限制sigma必须大于极小值,避免数值问题 bounds = ((-np.inf, 1e-4), (np.inf, np.inf)) params, cov = curve_fit(fit_lognormal, x, y_density, p0=p0, bounds=bounds) print(f"拟合参数:mu={params[0]:.4f}, sigma={params[1]:.4f}")
方案2:使用scipy.stats内置的fit方法(推荐)
scipy.stats.lognorm自带的fit方法基于极大似然估计,直接处理原始数据,无需手动处理直方图,稳定性更强:
from scipy.stats import lognorm # 直接拟合原始数据,固定loc=0(对数正态分布通常无偏移) sigma_fit, _, scale_fit = lognorm.fit(z_data, floc=0) mu_fit = np.log(scale_fit) print(f"拟合参数:mu={mu_fit:.4f}, sigma={sigma_fit:.4f}")
方案3:对数转换后拟合正态分布
对数正态分布的对数服从正态分布,转换后拟合正态分布更稳定:
from scipy.stats import norm import numpy as np # 确保数据全为正数 log_z = np.log(z_data[z_data > 0]) mu_fit, sigma_fit = norm.fit(log_z) print(f"拟合参数:mu={mu_fit:.4f}, sigma={sigma_fit:.4f}")
总结
优先选择方案2或方案3,这两种方法避开了直方图拟合的量级问题和数值不稳定风险,更适合对数正态分布的参数估计。
内容的提问来源于stack exchange,提问作者Roland
相关产品推荐
相关产品推荐

