You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

梯度下降算法定位最小值后损失函数激增问题求助

多元优化中梯度下降损失函数激增的问题与解决方案建议

我用梯度下降算法解决多元优化问题,算法整体表现尚可,但损失函数未呈现预期的单调下降趋势,在定位到最小值后会出现激增情况。我怀疑这与步长过大有关,但步长过小会导致算法运行缓慢,步长过大又会脱离解空间,陷入两难困境。

以下是我实现的优化器代码:

class spdg:
    def __init__(self, loss_function,
                 # a, c,
                 alpha_val,
                 gamma_val, max_iter, img_target, zernike,
                 momentum = 0.2,
                 cal_tolerance=1e-6):
        # Initialize gain parameters and decay factors
        # self.a = a
        # self.c = c
        self.alpha_val = alpha_val
        self.gamma_val = gamma_val

        self.loss = loss_function

        self.max_iter = max_iter
        self.A = max_iter / 10
        self.img_target = img_target
        self.zernike = zernike

        self.initial_entropy = image_entropy(self.img_target)
        self.momentum = momentum
        self.cal_tolerance = cal_tolerance

    def calc_loss(self, current_Aw):
        # print('current_Aw: %s' %str(current_Aw))

        Az = self.update_AZ(current_Aw)
        # print(Az[0][0])
        # print('Az: %s' %str(Az))
        """Evaluate the cost/loss function with a value of theta"""
        return self.loss(Az, self.img_target)

    def update_AZ(self, current_Aw):
        return zernike_plane(current_Aw, self.zernike)

    def minimise(self, current_Aw, optimizer_type='vanilla', verbose=False):
        k = 0  # initialize count
        cost_func_val = []
        Aw_values = []
        vk = 0

        previous_Aw = 0
        # while k < self.max_iter:

        while k < self.max_iter and \
                np.linalg.norm(previous_Aw - current_Aw) > self.cal_tolerance:

            previous_Aw = current_Aw
            cost_val = self.calc_loss(current_Aw)

            # get the current values for gain sequences
            # a_k = self.a / (k + 1 + self.A) ** self.alpha_val
            # c_k = self.c / (k + 1) ** self.gamma_val
            if verbose:
                print('iteration %d: %s with cost function value %.2f' % (k, str(current_Aw), cost_val))
            else:
                pass


            # get the random perturbation vector Bernoulli distribution with p=0.5
            delta = (np.random.randint(0, 2, current_Aw.shape) * 2 - 1) * self.gamma_val


            Aw_plus = current_Aw + delta
            Aw_minus = current_Aw - delta

            # measure the loss function at perturbations
            loss_plus = self.calc_loss(Aw_plus)
            loss_minus = self.calc_loss(Aw_minus)

            loss_delta = (loss_plus - loss_minus)
            # Aw_delta = Aw_plus - Aw_minus

            # compute the estimate of the gradient
            g_hat = loss_delta * delta

            # # update the estimate of the parameter
            if optimizer_type == 'vanilla':
                current_Aw = current_Aw - self.alpha_val * g_hat

            elif optimizer_type == 'momentum':

                vk_next = self.alpha_val * g_hat + self.momentum * vk
                current_Aw = current_Aw - vk_next
                vk = vk_next
            else:
                pass


            cost_val = self.calc_loss(current_Aw)
            cost_func_val.append(cost_val.squeeze())
            Aw_values.append(current_Aw)

            k += 1

        sol_idx = np.argmin(cost_func_val)
        Aw_estimate = Aw_values[sol_idx]
        print('optimal solution is found at %d iteration' % sol_idx)


        return Aw_estimate, cost_func_val

可行解决方案建议

  • 动态步长衰减:放弃固定的alpha_val,改用随迭代次数衰减的步长策略。例如alpha_k = self.alpha_val / np.sqrt(k + 1),迭代后期步长自动缩小,避免越过最优解。你代码中注释掉的a_k = self.a / (k + 1 + self.A) ** self.alpha_val就是类似思路,可恢复并调整参数测试。
  • 梯度裁剪:对估计得到的梯度g_hat进行范数限制,防止单次更新幅度过大。示例代码:
    max_grad_norm = 1.0  # 根据你的问题调整阈值
    norm = np.linalg.norm(g_hat)
    if norm > max_grad_norm:
        g_hat = g_hat * max_grad_norm / norm
    
  • Armijo线搜索:每次更新前通过线搜索确定最优步长,保证损失函数单调下降。核心逻辑是:先尝试当前步长,若损失上升则减半步长,直到满足loss(new_Aw) <= loss(old_Aw) + 0.1 * alpha * np.dot(g_hat.T, new_Aw - old_Aw)(0.1是可调整的松弛系数)。
  • 动量系数衰减:使用动量优化器时,随迭代次数减小动量系数,例如momentum_k = self.momentum * (1 - 1/(k+2)),避免后期动量导致冲过最优解。
  • 扰动幅度衰减:随机扰动的gamma_val也可随迭代衰减,比如gamma_k = self.gamma_val / np.sqrt(k+1),让梯度估计在后期更精确,减少更新方向的误差。
  • 提前停止机制:监控损失函数变化,当连续N次迭代损失上升或变化量小于阈值时终止迭代,避免在最小值附近震荡或激增。例如记录最近5次的损失值,若当前损失比历史最小值高出设定比例,则停止迭代。

内容的提问来源于stack exchange,提问作者Young Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:35:24