You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何针对收敛阶分析的有效数据段自动执行线性回归?

自动筛选收敛阶分析中的有效数据点执行线性回归

我正在绘制不同网格尺寸下的误差图,以研究模型的收敛阶。但实际场景中,"有效区间"并非始终存在:要么是网格过于粗糙(图1),要么误差已达到浮点精度极限(图2)。直接对全部数据执行线性回归无法得到正确结果,手动移除无效点后再用sklearn.linear_model执行线性回归的方案,因需处理约144个数据框,手动操作耗时过长。

我未在sklearn.linear_model中找到自动筛选有效数据点的功能,请问是否存在此类功能?或者有没有相关模块可以实现该需求?

我还设想了一种变通方法:计算所有连续点的斜率,保留对应"出现频率最高"的斜率的点来执行线性回归,但尚未实现。两幅图中,我用红色标出了希望用于线性回归的点。

绘图代码

import plotly.graph_objects as go
from io import StringIO
import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression
import plotly.express as px

df = pd.read_csv(StringIO("""hsize,errorl2
0.1,0.0126488001839
0.05,0.0130209787153
0.025,0.0111789092046
0.0125,0.00766945668181
0.00625,0.0012502455118
0.003125,0.000268314690697
"""))

def linear_regression_log(x, y):   # 对双对数数据执行线性回归,当前会用到所有数据
    keep = ~np.isnan(y) # 我的部分数据存在缺失情况
    log_x = np.log(np.array(x[keep]))
    log_y = np.log(np.array(y[keep]))

    model = LinearRegression()
    model.fit(log_x.reshape(-1, 1), log_y)

    slope = model.coef_[0]
    intercept = model.intercept_

    return slope, intercept

# 示例数据
x = df['hsize']
y = df['errorl2']
a, _ = linear_regression_log(x, y)

fig = go.Figure()
fig.add_trace(go.Scatter(x=x, y=y, mode='markers+lines'))
fig.update_layout(title=f"Slope: {a:.2f}")
fig.update_xaxes(type="log")
fig.update_yaxes(type="log")
fig.show()

绘图所用数据

  • 图1数据:
hsize,errorl2
0.1,0.0126488001839
0.05,0.0130209787153
0.025,0.0111789092046
0.0125,0.00766945668181
0.00625,0.0012502455118
0.003125,0.000268314690697
  • 图2数据:
hsize,errorl2
0.1,0.000713407986653
0.05,1.60793872143e-06
0.025,6.20078712336e-11
0.0125,2.99238475669e-13
0.00625,1.1731644955e-13
0.003125,1.88186766825e-13

内容的提问来源于stack exchange,提问作者Thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 11:32:47