You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让scipy.optimize.minimize(L-BFGS-B)执行完整maxiter迭代

问题描述

我想让scipy.optimize.minimize用L-BFGS-B方法跑完完整的20001次迭代,不被默认停止条件中断。当前代码只运行了约2500次就停止了,我已经尝试设置ftol: 0.0、gtol: 0.0并调大maxiter,但都没用。相关调用代码如下:

results = scipy.optimize.minimize(
    fun=get_loss_and_grads,
    x0=params,
    jac=True,
    method='L-BFGS-B',
    options={'maxiter': 20001, 'disp': True, 'ftol': 0.0, 'gtol': 0.0, 'maxcor': 50},
    callback=PrintLossCallback(print_freq=100)
)

完整代码如下:

def get_loss_and_grads(params):
    shapes = [w.shape for w in model.get_weights()]
    split_params = np.split(params, np.cumsum([np.prod(shape) for shape in shapes[:-1]]))
    weights = [tf.Variable(param.reshape(shape)) for param, shape in zip(split_params, shapes)]
    model.set_weights(weights)

    with tf.GradientTape() as tape:
        loss = combined_loss(X_train_tf[:, 0:1], X_train_tf[:, 1:2], x_init_tf, t_init_tf, u_initial_tf, x_bc_tf, t_bc_tf)
    gradients = tape.gradient(loss, model.trainable_variables)

    flattened_gradients = np.concatenate([grad.numpy().flatten() for grad in gradients])
    return loss.numpy().astype(np.float64), flattened_gradients

X_train_tf = tf.convert_to_tensor(X_train.astype(np.float32))
x_init_tf = tf.convert_to_tensor(x_init.astype(np.float32))
t_init_tf = tf.convert_to_tensor(t_init.astype(np.float32))
u_initial_tf = tf.convert_to_tensor(u_initial.astype(np.float32))
x_bc_tf = tf.convert_to_tensor(x_bc.astype(np.float32))
t_bc_tf = tf.convert_to_tensor(t_bc.astype(np.float32))


shapes = [w.shape for w in model.get_weights()]
params = np.concatenate([w.flatten() for w in model.get_weights()])

class PrintLossCallback:
    def __init__(self, print_freq=100):
        self.print_freq = print_freq
        self.iteration = 0

    def __call__(self, xk):
        self.iteration += 1
        if self.iteration % self.print_freq == 0:
            loss, _ = get_loss_and_grads(xk)
            print(f"Iteration {self.iteration}, Loss: {loss}")

start_time = time.time()


results = scipy.optimize.minimize(
    fun=get_loss_and_grads,
    x0=params,
    jac=True,
    method='L-BFGS-B',
    options={'maxiter': 20001, 'disp': True, 'ftol': 0.0, 'gtol': 0.0, 'maxcor': 50},
    callback=PrintLossCallback(print_freq=100)
)

optimized_params = results.x
split_params = np.split(optimized_params, np.cumsum([np.prod(shape) for shape in shapes[:-1]]))
model.set_weights([param.reshape(shape) for param, shape in zip(split_params, shapes)])

elapsed_time = time.time() - start_time
elapsed_minutes = elapsed_time / 60.0  # Convert seconds to minutes

print(f"Execution time: {elapsed_minutes:.2f} minutes")
解决方法

L-BFGS-B的停止条件不止ftol和gtol,你需要针对性调整以下参数:

  • 调大maxfun参数
    默认maxfun(函数调用次数上限)是15000,若迭代中函数调用次数触达这个值,优化会提前终止。建议将它设为maxiter的2倍以上,比如和你目标迭代次数匹配的40002,因为单次迭代可能触发多次函数调用。

  • 增大线搜索迭代次数maxls
    默认maxls是20,若线搜索多次失败,优化会停止。将其调大到100左右,能降低因线搜索触发的提前终止概率。

  • 排查停止原因
    运行完优化后,通过results.message和results.status可以明确停止原因:

    print("停止原因:", results.message)
    print("状态码:", results.status)
    

    常见状态码对应:

    • 0:收敛成功
    • 1:达到maxiter(正常跑完目标次数)
    • 2:达到maxfun(函数调用次数超限)
    • 3:线搜索失败
    • 4:梯度或函数值异常

修改后的options配置示例:

options={
    'maxiter': 20001,
    'disp': True,
    'ftol': 0.0,
    'gtol': 0.0,
    'maxcor': 50,
    'maxfun': 40002,
    'maxls': 100
}

另外,你的回调函数每次打印loss都会额外调用一次get_loss_and_grads,这会增加函数调用次数,可能提前触达maxfun限制。可以优化为在get_loss_and_grads中记录当前loss,避免重复计算:

# 新增全局变量记录当前loss
current_loss = None

def get_loss_and_grads(params):
    global current_loss
    shapes = [w.shape for w in model.get_weights()]
    split_params = np.split(params, np.cumsum([np.prod(shape) for shape in shapes[:-1]]))
    weights = [tf.Variable(param.reshape(shape)) for param, shape in zip(split_params, shapes)]
    model.set_weights(weights)

    with tf.GradientTape() as tape:
        loss = combined_loss(X_train_tf[:, 0:1], X_train_tf[:, 1:2], x_init_tf, t_init_tf, u_initial_tf, x_bc_tf, t_bc_tf)
    gradients = tape.gradient(loss, model.trainable_variables)

    flattened_gradients = np.concatenate([grad.numpy().flatten() for grad in gradients])
    current_loss = loss.numpy().astype(np.float64)
    return current_loss, flattened_gradients

# 修改回调函数,直接读取记录的loss
class PrintLossCallback:
    def __init__(self, print_freq=100):
        self.print_freq = print_freq
        self.iteration = 0

    def __call__(self, xk):
        self.iteration += 1
        if self.iteration % self.print_freq == 0:
            global current_loss
            print(f"Iteration {self.iteration}, Loss: {current_loss}")

内容的提问来源于stack exchange,提问作者Saif Ur Rehman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 19:53:16