Python曲线平滑、异常点剔除及非线性拟合技术咨询
异常点剔除+平滑拟合组合方案
先解决异常点问题,再做平滑和非线性拟合——不先剔除异常点,任何平滑/插值算法都会被异常点带偏,这是你之前方法没覆盖到的核心环节。
一、异常点剔除方法
1. 统计法(中位数绝对偏差MAD)
适合能明确数据波动范围的场景,比均值标准差更抗极端值:
import numpy as np def remove_outliers_mad(x, y, threshold=3): median_y = np.median(y) mad = np.median(np.abs(y - median_y)) if mad == 0: return x, y # 计算鲁棒Z分数 z_scores = 0.6745 * np.abs(y - median_y) / mad # 保留正常点 mask = z_scores < threshold return x[mask], y[mask]
2. 邻域拟合修正法
适合能判断局部趋势的情况,用周围点的拟合值修正异常点(或直接删除):
def remove_outliers_local_fit(x, y, window_size=5, threshold=2): cleaned_y = y.copy() for i in range(len(y)): # 取当前点的邻域窗口 start = max(0, i - window_size//2) end = min(len(y), i + window_size//2 + 1) window_x, window_y = x[start:end], y[start:end] # 排除当前点做局部线性拟合 idx = np.arange(len(window_x)) != (i - start) if np.sum(idx) < 2: continue coeffs = np.polyfit(window_x[idx], window_y[idx], 1) pred_y = np.polyval(coeffs, x[i]) # 残差超过阈值则替换为拟合值 if abs(y[i] - pred_y) > threshold * np.std(window_y[idx]): cleaned_y[i] = pred_y return x, cleaned_y
二、平滑处理(异常点剔除后再做)
用Savitzky-Golay滤波器,它基于局部多项式拟合,比高斯卷积更能保留曲线趋势:
from scipy.signal import savgol_filter # 先剔除异常点 cleaned_x, cleaned_y = remove_outliers_mad(original_x, original_y) # 平滑处理,window_length为奇数,polyorder为拟合阶数 smoothed_y = savgol_filter(cleaned_y, window_length=7, polyorder=2)
三、生成非线性方程
先根据曲线形态选模型(比如指数、幂函数、Sigmoid等),用非线性最小二乘拟合:
from scipy.optimize import curve_fit from sklearn.metrics import r2_score # 示例:指数模型,你可以换成符合你数据的模型 def nonlinear_model(x, a, b, c): return a * np.exp(b * x) + c # 拟合参数 params, _ = curve_fit(nonlinear_model, cleaned_x, cleaned_y) a, b, c = params # 验证拟合效果,R²越接近1越好 fitted_y = nonlinear_model(cleaned_x, a, b, c) r2 = r2_score(cleaned_y, fitted_y)
四、推荐流程
- 用MAD或邻域拟合方法剔除/修正异常点,得到干净数据集
- 对干净数据做Savitzky-Golay平滑(可选,看需求)
- 根据曲线趋势选择合适的非线性模型,拟合得到方程
内容的提问来源于stack exchange,提问作者Simon
相关产品推荐
相关产品推荐

