梯度下降算法定位最小值后损失函数激增问题求助
多元优化中梯度下降损失函数激增的问题与解决方案建议
我用梯度下降算法解决多元优化问题,算法整体表现尚可,但损失函数未呈现预期的单调下降趋势,在定位到最小值后会出现激增情况。我怀疑这与步长过大有关,但步长过小会导致算法运行缓慢,步长过大又会脱离解空间,陷入两难困境。
以下是我实现的优化器代码:
class spdg: def __init__(self, loss_function, # a, c, alpha_val, gamma_val, max_iter, img_target, zernike, momentum = 0.2, cal_tolerance=1e-6): # Initialize gain parameters and decay factors # self.a = a # self.c = c self.alpha_val = alpha_val self.gamma_val = gamma_val self.loss = loss_function self.max_iter = max_iter self.A = max_iter / 10 self.img_target = img_target self.zernike = zernike self.initial_entropy = image_entropy(self.img_target) self.momentum = momentum self.cal_tolerance = cal_tolerance def calc_loss(self, current_Aw): # print('current_Aw: %s' %str(current_Aw)) Az = self.update_AZ(current_Aw) # print(Az[0][0]) # print('Az: %s' %str(Az)) """Evaluate the cost/loss function with a value of theta""" return self.loss(Az, self.img_target) def update_AZ(self, current_Aw): return zernike_plane(current_Aw, self.zernike) def minimise(self, current_Aw, optimizer_type='vanilla', verbose=False): k = 0 # initialize count cost_func_val = [] Aw_values = [] vk = 0 previous_Aw = 0 # while k < self.max_iter: while k < self.max_iter and \ np.linalg.norm(previous_Aw - current_Aw) > self.cal_tolerance: previous_Aw = current_Aw cost_val = self.calc_loss(current_Aw) # get the current values for gain sequences # a_k = self.a / (k + 1 + self.A) ** self.alpha_val # c_k = self.c / (k + 1) ** self.gamma_val if verbose: print('iteration %d: %s with cost function value %.2f' % (k, str(current_Aw), cost_val)) else: pass # get the random perturbation vector Bernoulli distribution with p=0.5 delta = (np.random.randint(0, 2, current_Aw.shape) * 2 - 1) * self.gamma_val Aw_plus = current_Aw + delta Aw_minus = current_Aw - delta # measure the loss function at perturbations loss_plus = self.calc_loss(Aw_plus) loss_minus = self.calc_loss(Aw_minus) loss_delta = (loss_plus - loss_minus) # Aw_delta = Aw_plus - Aw_minus # compute the estimate of the gradient g_hat = loss_delta * delta # # update the estimate of the parameter if optimizer_type == 'vanilla': current_Aw = current_Aw - self.alpha_val * g_hat elif optimizer_type == 'momentum': vk_next = self.alpha_val * g_hat + self.momentum * vk current_Aw = current_Aw - vk_next vk = vk_next else: pass cost_val = self.calc_loss(current_Aw) cost_func_val.append(cost_val.squeeze()) Aw_values.append(current_Aw) k += 1 sol_idx = np.argmin(cost_func_val) Aw_estimate = Aw_values[sol_idx] print('optimal solution is found at %d iteration' % sol_idx) return Aw_estimate, cost_func_val
可行解决方案建议
- 动态步长衰减:放弃固定的
alpha_val,改用随迭代次数衰减的步长策略。例如alpha_k = self.alpha_val / np.sqrt(k + 1),迭代后期步长自动缩小,避免越过最优解。你代码中注释掉的a_k = self.a / (k + 1 + self.A) ** self.alpha_val就是类似思路,可恢复并调整参数测试。 - 梯度裁剪:对估计得到的梯度
g_hat进行范数限制,防止单次更新幅度过大。示例代码:max_grad_norm = 1.0 # 根据你的问题调整阈值 norm = np.linalg.norm(g_hat) if norm > max_grad_norm: g_hat = g_hat * max_grad_norm / norm - Armijo线搜索:每次更新前通过线搜索确定最优步长,保证损失函数单调下降。核心逻辑是:先尝试当前步长,若损失上升则减半步长,直到满足
loss(new_Aw) <= loss(old_Aw) + 0.1 * alpha * np.dot(g_hat.T, new_Aw - old_Aw)(0.1是可调整的松弛系数)。 - 动量系数衰减:使用动量优化器时,随迭代次数减小动量系数,例如
momentum_k = self.momentum * (1 - 1/(k+2)),避免后期动量导致冲过最优解。 - 扰动幅度衰减:随机扰动的
gamma_val也可随迭代衰减,比如gamma_k = self.gamma_val / np.sqrt(k+1),让梯度估计在后期更精确,减少更新方向的误差。 - 提前停止机制:监控损失函数变化,当连续N次迭代损失上升或变化量小于阈值时终止迭代,避免在最小值附近震荡或激增。例如记录最近5次的损失值,若当前损失比历史最小值高出设定比例,则停止迭代。
内容的提问来源于stack exchange,提问作者Young Wang
相关产品推荐
相关产品推荐

