Keras与Levenberg-Marquardt优化对接失败问题求助
我尝试将Keras与scipy.optimize.least_squares中的Levenberg-Marquardt优化方法结合,目标是从训练数据中拟合出y=2*x的关系。在model_residuals函数中,我分别用Keras模型和简单线性回归模型计算残差r和r2,两者几乎完全一致,但返回r2时模型能正常收敛到最优解,返回r时则完全不收敛。请问为何会出现这种差异?问题出在哪里?
完整代码如下:
import sys import numpy as np import matplotlib.pyplot as plt from keras.models import Sequential from keras.layers import Dense from keras.optimizers import * from scipy.optimize import least_squares from copy import deepcopy # Data generation. np.random.seed(0); n_samples = 20; X = np.linspace(-1, 1, n_samples) y = 2 * X + np.random.randn(n_samples) * 0.1; # The equation / model we want to retrieve is Y = 2X. # Construction of the Keras ANN model. model = Sequential() model.add(Dense(1, input_dim=1, use_bias=False, activation='linear')) ########################################### # Function which returns the residuals as a list. def model_residuals(params, x, y): print(" params0 = ",params); weights = deepcopy(params); # In order not to change "params". weights = weights.reshape((1, 1)) model.layers[0].set_weights([weights]); # Set the weights of the model. current_params = model.get_weights()[0].flatten() print(" weights in model : ",params); r = model.predict(x).flatten() - y; # Residuals computed using the Keras model. r2 = params[0]*x-y; # Residuals computed using a simple linear regression model. print(" r ",r); print(" r2 ",r2); print("\n","\n","\n"); # The residual list r and r2 are always almost the same for any parameter value! squared_error = np.mean(r**2); print("Squared Error:", squared_error); squared_error2 = np.mean(r2**2); print("Squared Error2:", squared_error2); # Likewise, the squared errors are almost always the same for any parameter value! return r ######################################## init_params = model.get_weights()[0].flatten() # Weights 1D-Array init_params = [-9]; # Let us always initialise the coefficient of the simple linear model by the value -2 as it is more challenging for the ANN to retrieve it. print(" init_params = ",init_params); max_nfev = 1000 # Maximal number of iterations. result = least_squares(model_residuals, init_params, args=(X, y), method='lm', xtol=1e-6, max_nfev = max_nfev, jac='2-point'); # Fitting the coefficient to the data using either the simple linear regression model (in which case the function "model_residuals" must return 'r2') or the ANN from Keras (in which case the function "model_residuals" must return 'r') print("\n","\n"," result = ",result,"\n","\n"); # Set the optimised weights. opt_weights = [result.x.reshape((1, 1))] print("\n","\n"," opt_weights = ",opt_weights,"\n","\n"); model.layers[0].set_weights(opt_weights); # Set the optimised weights into the model. # Prediction with the optimised model. X_test = np.linspace(-1, 1, 40); y_test = 2 * X_test; y_pred = model.predict(X_test); # Plot the quality of the fit. plt.scatter(X_test, y_test, color='blue', label='Data') plt.plot(X_test, y_pred.flatten(), color='red', label='Fit') plt.xlabel('X') plt.ylabel('y') plt.legend() plt.show()
核心原因:Keras的predict方法干扰了scipy的数值导数计算
scipy的least_squares使用jac='2-point'时,会通过微小扰动参数并计算残差变化来数值近似雅可比矩阵。但Keras的predict方法在后台会构建计算图(即使是简单线性层),还可能对权重进行隐式的类型转换或状态缓存,导致scipy在计算数值导数时得到错误的残差变化量,最终让LM算法的优化方向完全错误,无法收敛。
具体差异:
- 用
params[0]*x-y计算r2时,是纯numpy操作,无额外状态,scipy的数值扰动能准确捕捉残差随参数的变化,雅可比矩阵计算正确,优化正常收敛。 - 用Keras的
model.predict(x)计算r时,虽然单次残差结果和r2几乎一致,但Keras的内部处理会破坏scipy数值导数计算的准确性——比如权重被转为TensorFlow张量后,后续扰动时的状态重置出现偏差,导致残差变化量失真,LM算法找不到正确的优化方向。
解决方案:用numpy手动计算模型输出,绕开Keras的predict
既然模型是无偏置的线性层,完全可以不用调用Keras的predict,直接用numpy计算权重与输入的乘积,既保留Keras模型结构,又避免计算图的干扰:
修改model_residuals函数中的残差计算部分:
# 替换原有的r = model.predict(x).flatten() - y r = (params[0] * x) - y # 同时保留模型权重更新,方便后续用Keras做预测 model.layers[0].set_weights([params.reshape(1,1)])
这样修改后,残差计算是纯numpy操作,scipy的数值导数计算能正常工作,LM算法可以收敛到最优解,同时Keras模型的权重也会被正确更新。
额外验证方法
可以在model_residuals中手动计算残差对参数的导数(对于线性模型,导数就是输入x),然后和scipy计算的雅可比矩阵对比。当使用Keras的predict时,scipy得到的雅可比矩阵会和真实值有偏差,这就是不收敛的直接证据。
内容的提问来源于stack exchange,提问作者Marc Fischer

