在Python中使用LOESS/LOWESS时,如何为数据点设置权重?
在Statsmodels的LOWESS中使用自定义数据点权重
Statsmodels库的sm.nonparametric.lowess函数支持通过weights参数传入自定义数据点权重,直接实现你的需求。
修改后的示例代码
import numpy as np import statsmodels.api as sm import matplotlib.pyplot as plt # 创建数据 x = np.random.uniform(low=-2*np.pi, high=2*np.pi, size=500) y = np.sin(x) + np.random.normal(size=len(x)) # 创建自定义权重(1到4的整数,模拟不同数据点的重要性) weights = np.random.randint(1, 5, size=500) # 带权重的LOWESS拟合 lowess = sm.nonparametric.lowess # 传入weights参数,指定每个数据点的权重 z_weighted = lowess(y, x, weights=weights) z_unweighted = lowess(y, x) # 无权重版本用于对比 # 绘图对比 plt.figure(figsize=(12, 6)) # 用点的大小可视化权重:权重越大,点越大 plt.scatter(x, y, s=weights*10, label='Data (size = weight)', alpha=0.5) plt.plot(z_unweighted[:, 0], z_unweighted[:, 1], label='LOWESS (unweighted)', color='red') plt.plot(z_weighted[:, 0], z_weighted[:, 1], label='LOWESS (weighted)', color='blue', linestyle='--') plt.legend() plt.show()
关键说明
weights参数接受一个与输入数据长度相同的数组,每个元素对应一个数据点的权重值,数值越大表示该点在局部回归中的优先级越高。- 示例中用点的大小直观展示权重差异,你可以根据实际业务需求调整权重的生成逻辑。
内容的提问来源于stack exchange,提问作者jss367
相关产品推荐
相关产品推荐

