scipy.optimize.curve_fit中sigma参数含义及仪器不确定度纳入拟合的方法咨询
scipy.optimize.curve_fit中sigma参数含义及仪器不确定度纳入拟合的方法咨询
假设我有一个通过重复采样测量某个量的简单模型。根据统计学知识,均值的误差可以通过以下公式计算:
$$u = \frac{\hat\sigma}{\sqrt{n}}$$
其中
$$\hat\sigma = \frac{\sum(x-\bar{x})}{\sqrt{n-1}}$$
当采样次数 $n \rightarrow \infty$ 时,这种不确定度会消失,但仪器本身还有B类不确定度。通常我们认为这两种不确定度可以通过方和根的方式叠加。
但如果是通过曲线拟合来处理数据呢?我尝试做了类似这样的拟合:

这可以用scipy.optimize.curve_fit轻松实现,代码如下:
import numpy as np from scipy.optimize import curve_fit import matplotlib.pyplot as plt data_length = 10000 x = np.array(range(data_length)) y = 5 * np.ones(data_length) + np.random.normal(0, 1, data_length) u_y = 1 * np.ones(data_length) def func(x, b): return x*0 + b popt, pcov = curve_fit(func, x,y, sigma=u_y, absolute_sigma=False) # 生成绘图 plt.errorbar(x, y, yerr=u_y, fmt='o') plt.plot(x, func(x, *popt), 'r-') plt.show() print(popt) print(np.sqrt(np.diag(pcov)))
结果发现当 $n \rightarrow \infty$ 时,拟合参数的误差趋近于0,这意味着可以通过大量测量来突破仪器的精度限制,这显然不符合实际。另外我发现如果把sigma缩放为sigma=1e6*u_y,拟合参数的误差并没有变化,这说明这个函数里的sigma似乎并没有对应仪器给出的测量不确定度。
我查阅了文档,发现应该设置absolute_sigma=True。我尝试了这个参数,但它似乎总是返回1/np.sqrt(data_length),而且数据集本身的随机性并没有起到作用。另外,当 $n \rightarrow \infty$ 时误差依然会消失,这让我很困惑。
我想知道:
- 曲线拟合中正确纳入仪器误差的方式是什么?
- 这里的
sigma参数到底代表什么含义? - 在有多个参数的真实拟合中,该如何纳入这个仪器不确定度项?
我还尝试了线性拟合的情况:
sampling_points = int(1e2) x = np.linspace(0, 1, sampling_points) y = 5 * x + np.random.normal(0, 1, sampling_points) u_y = 1 * np.ones(sampling_points) def func(x, a, b): return x*a + b popt, pcov = curve_fit(func, x,y, sigma=1*u_y, absolute_sigma=False) print("sample points: ", sampling_points) print("k={pop[0]:.2f}, b={pop[1]:.2f}".format(pop=popt)) print("error of k: ", np.sqrt(pcov[0,0])) print("error of b: ", np.sqrt(pcov[1,1]))
结果依然是增加数据点可以完全消除误差:
当采样点为100时:
sample points: 100 k=4.74, b=0.23 error of k: 0.37312467183800596 error of b: 0.2159669523150594
当采样点为1e6时:
sample points: 1000000 k=5.01, b=-0.00 error of k: 0.0034627280798220374 error of b: 0.0019992075688566534
备注:内容来源于stack exchange,提问作者YuanFeng Sheh
相关产品推荐
相关产品推荐

