You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单变量二元场景下loess与lowess差异及结果匹配方法问询

loess与lowess的差异及匹配方法

问题背景

首先生成模拟数据:

probs <- seq(0,0.3, by=0.001)

targets <- sapply(probs, function(p){
  sample(c(0,1),size=1,prob = c(1-p,p))
})

使用loess对targets做平滑并绘图(作为概率估计):

require(magrittr)
loess(targets~probs,span=0.3) %>% predict %>% {plot(. ~ probs)}

但无论调整f值,lowess都无法得到相同效果:

lowess(x = probs, y = targets, f = 0.01) %>% with(plot(y ~ x))

核心问题:两者结果为何存在差异?能否让lowess输出与loess匹配的结果?
注:需理解差异,并在仅提供lowess的Python statsmodels中应用相关结果。

两者差异原因

  • 拟合阶数不同:loess默认采用二次多项式做局部拟合,而lowess默认使用线性多项式,这导致前者的平滑曲线更贴合非线性趋势,后者更容易出现波动。
  • 参数细节差异:loess的span与lowess的f虽都是控制局部窗口的样本比例,但loess对权重的计算逻辑更精细化,尤其是在处理边界点时的调整策略不同。
  • 离散数据适配性:你的targets是0-1离散值,二次拟合的loess对这类数据的平滑效果更稳定,而线性拟合的lowess对离散点的拟合精度天生更弱。

让lowess匹配loess结果的方法

可以通过调整lowess的参数逼近loess的效果:

  1. 提升拟合阶数:设置degree=2,对应loess默认的二次多项式拟合。
  2. 对齐窗口比例:将f参数设为与loess的span相同的值(示例中设为f=0.3)。
  3. 权重类型统一:两者默认均使用tricube权重,无需额外调整。

调整后的lowess代码示例:

lowess(x = probs, y = targets, f = 0.3, degree = 2) %>% with(plot(y ~ x))

在Python statsmodels中的应用

statsmodels中的lowess函数(statsmodels.nonparametric.smoothers_lowess.lowess)可通过以下参数调整逼近R中loess的效果:

  • 设置frac=0.3(对应R中loess的span=0.3)
  • 开启is_sorted=True提升效率(因probs是有序序列)

示例代码:

import numpy as np
from statsmodels.nonparametric.smoothers_lowess import lowess

# 生成模拟数据
probs = np.arange(0, 0.3, 0.001)
targets = np.array([np.random.choice([0, 1], p=[1-p, p]) for p in probs])

# 逼近loess效果的lowess平滑
smoothed_result = lowess(targets, probs, frac=0.3, is_sorted=True)
# 提取平滑后的值
smoothed_y = smoothed_result[:, 1]

若需要严格的二次局部拟合,可结合statsmodels的局部多项式回归工具,或自定义局部加权二次回归逻辑来实现。

内容的提问来源于stack exchange,提问作者runr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 13:09:53