You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras与Levenberg-Marquardt优化对接失败问题求助

问题:Keras模型结合scipy.least_squares(LM方法)不收敛,手动计算残差却正常收敛的原因

我尝试将Keras与scipy.optimize.least_squares中的Levenberg-Marquardt优化方法结合,目标是从训练数据中拟合出y=2*x的关系。在model_residuals函数中,我分别用Keras模型和简单线性回归模型计算残差r和r2,两者几乎完全一致,但返回r2时模型能正常收敛到最优解,返回r时则完全不收敛。请问为何会出现这种差异?问题出在哪里?

完整代码如下:

import sys 
import numpy as np
import matplotlib.pyplot as plt
from keras.models import Sequential
from keras.layers import Dense
from keras.optimizers import *
from scipy.optimize import least_squares
from copy import deepcopy


# Data generation.
np.random.seed(0); n_samples = 20; 
X = np.linspace(-1, 1, n_samples)
y = 2 * X + np.random.randn(n_samples) * 0.1; # The equation / model we want to retrieve is Y = 2X. 

# Construction of the Keras ANN model. 
model = Sequential()
model.add(Dense(1, input_dim=1, use_bias=False, activation='linear'))


###########################################
# Function which returns the residuals as a list.
def model_residuals(params, x, y):
    print(" params0 = ",params);
    weights = deepcopy(params); # In order not to change "params". 
    weights = weights.reshape((1, 1))
    model.layers[0].set_weights([weights]); # Set the weights of the model. 
    
    current_params = model.get_weights()[0].flatten()
    print(" weights in model : ",params);  
    
    
    r = model.predict(x).flatten() - y; # Residuals computed using the Keras model.
    r2 = params[0]*x-y; # Residuals computed using a simple linear regression model.
    
        
    print(" r ",r); print(" r2 ",r2); print("\n","\n","\n"); 
    # The residual list r and r2 are always almost the same for any parameter value! 

    squared_error = np.mean(r**2); 
    print("Squared Error:", squared_error); 
    squared_error2 = np.mean(r2**2); 
    print("Squared Error2:", squared_error2); 
    
    # Likewise, the squared errors are almost always the same for any parameter value! 
    
    return r
########################################


init_params = model.get_weights()[0].flatten()  # Weights 1D-Array
init_params = [-9]; # Let us always initialise the coefficient of the simple linear model by the value -2 as it is more challenging for the ANN to retrieve it. 
print(" init_params = ",init_params); 


max_nfev = 1000  # Maximal number of iterations. 

result = least_squares(model_residuals, init_params, args=(X, y), method='lm', xtol=1e-6, max_nfev = max_nfev, jac='2-point'); # Fitting the coefficient to the data using either the simple linear regression model (in which case the function "model_residuals" must return 'r2') or the ANN from Keras (in which case the function "model_residuals" must return 'r')

print("\n","\n"," result = ",result,"\n","\n"); 

# Set the optimised weights. 
opt_weights = [result.x.reshape((1, 1))] 

print("\n","\n"," opt_weights = ",opt_weights,"\n","\n"); 

model.layers[0].set_weights(opt_weights); # Set the optimised weights into the model.


# Prediction with the optimised model. 
X_test = np.linspace(-1, 1, 40); 
y_test = 2 * X_test; 
y_pred = model.predict(X_test); 


# Plot the quality of the fit. 
plt.scatter(X_test, y_test, color='blue', label='Data')
plt.plot(X_test, y_pred.flatten(), color='red', label='Fit')
plt.xlabel('X')
plt.ylabel('y')
plt.legend()
plt.show()

核心原因:Keras的predict方法干扰了scipy的数值导数计算

scipy的least_squares使用jac='2-point'时,会通过微小扰动参数并计算残差变化来数值近似雅可比矩阵。但Keras的predict方法在后台会构建计算图(即使是简单线性层),还可能对权重进行隐式的类型转换或状态缓存,导致scipy在计算数值导数时得到错误的残差变化量,最终让LM算法的优化方向完全错误,无法收敛。

具体差异:

  • 用params[0]*x-y计算r2时,是纯numpy操作,无额外状态,scipy的数值扰动能准确捕捉残差随参数的变化,雅可比矩阵计算正确,优化正常收敛。
  • 用Keras的model.predict(x)计算r时,虽然单次残差结果和r2几乎一致,但Keras的内部处理会破坏scipy数值导数计算的准确性——比如权重被转为TensorFlow张量后,后续扰动时的状态重置出现偏差,导致残差变化量失真,LM算法找不到正确的优化方向。

解决方案:用numpy手动计算模型输出,绕开Keras的predict

既然模型是无偏置的线性层,完全可以不用调用Keras的predict,直接用numpy计算权重与输入的乘积,既保留Keras模型结构,又避免计算图的干扰:

修改model_residuals函数中的残差计算部分:

# 替换原有的r = model.predict(x).flatten() - y
r = (params[0] * x) - y
# 同时保留模型权重更新,方便后续用Keras做预测
model.layers[0].set_weights([params.reshape(1,1)])

这样修改后,残差计算是纯numpy操作,scipy的数值导数计算能正常工作,LM算法可以收敛到最优解,同时Keras模型的权重也会被正确更新。

额外验证方法

可以在model_residuals中手动计算残差对参数的导数(对于线性模型,导数就是输入x),然后和scipy计算的雅可比矩阵对比。当使用Keras的predict时,scipy得到的雅可比矩阵会和真实值有偏差,这就是不收敛的直接证据。

内容的提问来源于stack exchange,提问作者Marc Fischer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 03:37:25