如何检测线性数据中的Hump?技术实现方案问询
识别线性分布数据中的凸起区间数据点
我有一组x、y数据,整体呈线性分布,但某一区间的数据偏离直线形成凸起(Hump)。目标是识别出这个凸起内的所有数据点。
我尝试用全部数据拟合线性回归趋势线,计算残差后剔除残差较大的区间,但效果不佳——拟合的趋势线未贴合正确的线性点位。以下是我的原代码及运行结果:
原代码
import numpy as np import matplotlib.pyplot as plt from sklearn.linear_model import LinearRegression # Data provided x = np.array([134, 147, 161, 175, 190, 206, 222, 237, 251, 263, 275, 291, 300, 312, 324, 337, 349, 360, 372, 382]).reshape(-1, 1) y = np.array([0.788875116, 0.692846919, 0.605305046, 0.738780558, 0.826074803, 0.871572936, 0.776701184, 0.646403726, 0.677606953, 0.615950052, 0.357934847, 0.267171728, 0.217483944, 0.155336037, 0.071882007, 0.029383778, -0.008773924, -0.050609993, -0.102372909, -0.148741651]) # Fit linear regression model = LinearRegression() model.fit(x, y) y_pred = model.predict(x) # Calculate residuals residuals = y - y_pred std_dev = np.std(residuals) # Identify hump points by adjusting alpha alpha = 2 / 3 threshold = std_dev * alpha hump_points = np.where(np.abs(residuals) > threshold)[0] # Visualize plt.figure(figsize=(12, 6)) plt.plot(x, y, 'o-', label='Data', markersize=6) plt.plot(x, y_pred, 'r--', label='Fitted Line', linewidth=2) plt.axhline(0, color='gray', linestyle='--', label='Zero Residual') plt.scatter(x[hump_points], y[hump_points], color='orange', label='Hump Point', s=100) plt.legend() plt.xlabel('X') plt.ylabel('Y') plt.title('Detecting Hump in Data') plt.grid(True) plt.show() # Print hump points hump_info = [(x[point][0], y[point]) for point in hump_points]
原代码运行结果

问题分析
原方法的核心缺陷是用包含凸起的全部数据拟合直线,导致拟合线被凸起数据拉偏,残差计算失去了正确的参考基准,自然无法准确识别凸起点。
改进方案:稳健迭代拟合+残差筛选
通过逐步移除残差过大的点,让拟合线逐渐贴合真实的线性趋势,再基于这条正确的趋势线筛选凸起点:
改进后代码
import numpy as np import matplotlib.pyplot as plt from sklearn.linear_model import LinearRegression # 原始数据 x = np.array([134, 147, 161, 175, 190, 206, 222, 237, 251, 263, 275, 291, 300, 312, 324, 337, 349, 360, 372, 382]).reshape(-1, 1) y = np.array([0.788875116, 0.692846919, 0.605305046, 0.738780558, 0.826074803, 0.871572936, 0.776701184, 0.646403726, 0.677606953, 0.615950052, 0.357934847, 0.267171728, 0.217483944, 0.155336037, 0.071882007, 0.029383778, -0.008773924, -0.050609993, -0.102372909, -0.148741651]) # 迭代拟合函数:逐步移除残差过大的点,得到贴合线性部分的稳健模型 def get_robust_fit(x, y, threshold_factor=2, max_iter=3): current_x, current_y = x.copy(), y.copy() for _ in range(max_iter): model = LinearRegression() model.fit(current_x, current_y) y_pred = model.predict(current_x) residuals = current_y - y_pred std_dev = np.std(residuals) # 保留残差在阈值内的点 mask = np.abs(residuals) <= threshold_factor * std_dev current_x, current_y = current_x[mask], current_y[mask] # 用最终筛选后的点重新拟合 final_model = LinearRegression() final_model.fit(current_x, current_y) return final_model # 获取稳健拟合的线性模型 robust_model = get_robust_fit(x, y) y_robust_pred = robust_model.predict(x) # 基于稳健拟合计算残差,筛选向上偏离的凸起点 residuals_robust = y - y_robust_pred std_robust = np.std(residuals_robust) # 针对向上凸起,只筛选正残差超过阈值的点(阈值可根据数据调整) hump_threshold = 1.5 * std_robust hump_points = np.where(residuals_robust > hump_threshold)[0] # 可视化结果 plt.figure(figsize=(12, 6)) plt.plot(x, y, 'o-', label='原始数据', markersize=6) plt.plot(x, y_robust_pred, 'r--', label='稳健拟合直线', linewidth=2) plt.scatter(x[hump_points], y[hump_points], color='orange', label='凸起数据点', s=100, zorder=5) plt.legend() plt.xlabel('X') plt.ylabel('Y') plt.title('识别线性数据中的凸起区间') plt.grid(True) plt.show() # 输出凸起点信息 hump_info = [(x[point][0], y[point]) for point in hump_points] print("识别出的凸起点:", hump_info)
方案说明
- 迭代稳健拟合:通过3轮迭代,每次移除残差超过2倍标准差的点,最终用筛选后的纯线性数据拟合趋势线,确保趋势线贴合真实线性部分。
- 针对性残差筛选:因为凸起是向上偏离,所以只筛选正残差超过1.5倍标准差的点,避免误判向下的正常数据。
- 参数可调:
threshold_factor(迭代剔除阈值)和hump_threshold(凸起判定阈值)可根据数据分布灵活调整。
内容的提问来源于stack exchange,提问作者Luks
相关产品推荐
相关产品推荐

