Scipy curve_fit多系列数据拟合:为多条曲线求解统一拟合参数
嘿,这个多区域共用参数拟合的问题我之前也碰到过,直接拼接数据确实容易出问题——毕竟不同区域的数据可能量级、分布差异很大,拟合的时候很容易偏向某一块区域,导致整体效果拉胯。给你几个可行的解决方案,亲测有效:
解决方案1:用自定义全局损失函数 +
scipy.optimize.minimize替代curve_fit curve_fit默认是对所有输入点计算均方误差,但它没法灵活处理多区域的联合拟合需求。换成minimize的话,你可以自定义一个全局损失函数,把每个区域的误差加起来,让优化器同时考虑所有区域的拟合效果。
举个代码例子:
import numpy as np from scipy.optimize import minimize # 先把所有区域的数据整理成列表,每个元素是(x, 橙色曲线y值, 目标y值) regions = [ (x_region1, orange_y_region1, target_y_region1), (x_region2, orange_y_region2, target_y_region2), # 有多少个区域就加多少个 ] # 这里是你的非线性函数,替换成你实际的实现 def non_linear_function(x, c1, c2, c3): return c1 * np.exp(-c2 * x) + c3 * np.power(x, 2) # 定义全局损失函数:所有区域的均方误差之和 def total_loss(params): c1, c2, c3 = params total_error = 0.0 for x, orange_y, target_y in regions: # 计算当前参数下的修正后曲线 damage = non_linear_function(x, c1, c2, c3) predicted_y = np.multiply(damage, orange_y) # 累加当前区域的误差(这里用均方误差,你也可以换成其他误差指标,比如绝对误差) total_error += np.sum(np.square(predicted_y - target_y)) return total_error # 初始参数猜测:可以用单个区域拟合出来的参数当初始值,加快收敛 initial_guess = [1.0, 0.5, 0.1] # 换成你自己的初始值 # 调用优化器求解最优参数,选一个适合你问题的优化方法,L-BFGS-B适合带约束的情况,没有约束也可以用Nelder-Mead result = minimize(total_loss, initial_guess, method='L-BFGS-B') # 得到最优参数 best_c1, best_c2, best_c3 = result.x
解决方案2:给不同区域设置权重(可选)
如果某些区域的拟合优先级更高,你可以在损失函数里给对应的区域加上权重系数,让优化器更重视这些区域的误差:
# 给每个区域分配权重,比如区域1权重0.7,区域2权重0.3,总和不一定是1,按需求调整 region_weights = [0.7, 0.3] def weighted_total_loss(params): c1, c2, c3 = params total_error = 0.0 for (x, orange_y, target_y), weight in zip(regions, region_weights): damage = non_linear_function(x, c1, c2, c3) predicted_y = np.multiply(damage, orange_y) # 乘以权重后累加误差 total_error += weight * np.sum(np.square(predicted_y - target_y)) return total_error # 同样调用minimize求解 result = minimize(weighted_total_loss, initial_guess, method='L-BFGS-B')
解决方案3:标准化数据(解决量级差异问题)
如果不同区域的y值量级差得很大,拟合时会自动偏向量级大的区域,这时候可以先对每个区域的数据做标准化(比如归一化到[0,1]区间),拟合完之后再还原回原始量级:
# 先标准化每个区域的数据,同时保存缩放参数用于还原 normalized_regions = [] scalers = [] for x, orange_y, target_y in regions: # 对目标y和橙色曲线y做最小-最大归一化 target_min, target_max = target_y.min(), target_y.max() orange_min, orange_max = orange_y.min(), orange_y.max() norm_target = (target_y - target_min) / (target_max - target_min) norm_orange = (orange_y - orange_min) / (orange_max - orange_min) normalized_regions.append((x, norm_orange, norm_target)) scalers.append((target_min, target_max, orange_min, orange_max)) # 用标准化后的数据定义损失函数 def normalized_loss(params): c1, c2, c3 = params total_error = 0.0 for x, norm_orange, norm_target in normalized_regions: damage = non_linear_function(x, c1, c2, c3) predicted_y = np.multiply(damage, norm_orange) total_error += np.sum(np.square(predicted_y - norm_target)) return total_error result = minimize(normalized_loss, initial_guess, method='L-BFGS-B') best_c1, best_c2, best_c3 = result.x # 如果需要还原到原始量级,可以用这个函数 def predict_original(x, orange_y, scaler): target_min, target_max, orange_min, orange_max = scaler norm_orange = (orange_y - orange_min) / (orange_max - orange_min) damage = non_linear_function(x, best_c1, best_c2, best_c3) norm_pred = np.multiply(damage, norm_orange) # 还原回原始范围 return norm_pred * (target_max - target_min) + target_min
为什么直接拼接数据不行?
当你把所有区域的数据拼接成一维数组时,curve_fit会把每个点当成同等重要的样本,但不同区域的数据可能在量级、点的数量上差异很大——比如某个区域有1000个点,另一个只有100个,或者某个区域的y值是另一个的100倍,拟合结果就会被占优的区域牵着走,其他区域的拟合效果自然就差了。而上面的方法可以让你更灵活地控制每个区域的贡献,从而得到全局最优的参数。
内容的提问来源于stack exchange,提问作者Polygondwana
相关产品推荐
相关产品推荐

