如何用Python为具有特殊分布的数据集绘制拟合(优化)函数?
自定义分布拟合与绘制方案
1. 手动构造匹配形态的复合模型
你的数据分布不属于标准统计分布,需要手动拼接函数来覆盖所有特征:
- x<30区间:用指数项模拟初始陡降,比如
a*np.exp(-b*x) - 30<x<80区间:加入波动项(正弦/余弦或多项式)来拟合小峰谷,比如
c*np.sin(d*x + e) - x>80区间:用快速衰减项模拟骤降,比如
f*np.exp(-g*(x-80))
推荐平滑过渡的连续模型(避免分段的生硬断点):
import numpy as np def custom_smooth_model(x, a, b, c, d, e, f, g, h): # 用sigmoid实现平滑分段过渡 transition_30 = 1 / (1 + np.exp(h*(x-30))) # x=30时过渡到中间段 transition_80 = 1 / (1 + np.exp(h*(x-80))) # x=80时过渡到尾段 # 各段组合 base_decay = a * np.exp(-b*x) mid_wave = c * np.sin(d*x + e) tail_drop = f * np.exp(-g*(x-80)) return base_decay*transition_30 + (base_decay + mid_wave)*(1-transition_30)*transition_80 + tail_drop*(1-transition_80)
2. 用curve_fit优化参数
放弃Fitter(仅适配标准分布),改用scipy.optimize.curve_fit来拟合自定义模型:
from scipy.optimize import curve_fit import matplotlib.pyplot as plt # 加载你的数据(替换为实际数据加载方式) data = np.loadtxt("your_dataset.txt") # 生成直方图密度数据(用于拟合参考) counts, bins = np.histogram(data, bins=50, density=True) bin_centers = (bins[:-1] + bins[1:])/2 # 取bin中点作为拟合x值 # 手动输入初始参数猜测(根据直方图估算:峰值、衰减率、波动幅度等) initial_guess = [max(counts), 0.1, 0.05, 0.1, 0, 0.01, 0.2, 1] # 执行拟合 popt, pcov = curve_fit(custom_smooth_model, bin_centers, counts, p0=initial_guess)
3. 绘制拟合结果
把原始直方图和拟合曲线可视化:
# 生成拟合曲线的连续x值 x_fit = np.linspace(min(data), max(data), 1000) y_fit = custom_smooth_model(x_fit, *popt) # 绘图 plt.hist(data, bins=50, density=True, alpha=0.5, label="原始数据直方图") plt.plot(x_fit, y_fit, 'r-', linewidth=2, label="拟合曲线") plt.xlabel("x") plt.ylabel("密度") plt.legend() plt.show()
4. 调优技巧
- 如果30-80区间的峰谷是非周期性的,把正弦项换成二次多项式(
c*x² + d*x + e) - 若x>80处几乎无数据,可直接删除尾段项,简化模型
- 初始参数猜测要贴近实际:比如
a设为直方图峰值,b设为能匹配初始衰减速度的数值,避免拟合陷入局部最优 - 如需后续统计推断,可把模型封装为
scipy.stats.rv_continuous自定义分布类:from scipy.stats import rv_continuous class CustomDist(rv_continuous): def _pdf(self, x, a, b, c, d, e, f, g, h): return custom_smooth_model(x, a, b, c, d, e, f, g, h) custom_dist = CustomDist() fitted_params = custom_dist.fit(data, p0=initial_guess)
内容的提问来源于stack exchange,提问作者J. Chae
相关产品推荐
相关产品推荐

