TensorFlow自定义loss function问题:拟合指数函数为何收敛至常数?
问题描述
第一段代码可在[0,1]区间用神经网络近似指数函数exp(x),运行后能正常收敛:
# code works import tensorflow as tf import numpy as np import matplotlib.pyplot as plt # fit an exponential n = 101 x = np.linspace(start=0, stop=1, num=n) y_e = np.exp(x) # any odd neural net with sufficient degrees of freedom model = tf.keras.models.Sequential([ tf.keras.layers.Dense(units=1, input_shape=[1]), tf.keras.layers.Dense(units=50, activation="softmax"), tf.keras.layers.Dense(units=1) ]) # loss function def loss(y_true, y_pred): L_ode = tf.reduce_mean(tf.square(y_pred - y_true), axis=-1) return L_ode model.compile('adam', loss) model.fit(x, y_e, epochs=100, batch_size=1) y_NN = model.predict(x).flatten() plt.plot(x, y_NN, color='blue') plt.plot(x, y_e, color='red') plt.title('NN (blue) and exp (red)')
修改loss函数并在model.fit中传入tf.zeros(n)作为y_true后,代码运行正常但模型收敛至常数,无法近似exp(x):
import tensorflow as tf import numpy as np import matplotlib.pyplot as plt # fit an exponential n = 101 x = np.linspace(start=0, stop=1, num=n) y_e = np.exp(x) # any odd neural net with sufficient degrees of freedom model = tf.keras.models.Sequential([ tf.keras.layers.Dense(units=1, input_shape=[1]), tf.keras.layers.Dense(units=50, activation="softmax"), tf.keras.layers.Dense(units=1) ]) # loss function def loss(y_true, y_pred): L_ode = tf.reduce_mean(tf.square((y_pred - y_e) - y_true), axis=-1) return L_ode model.compile('adam', loss) model.fit(x, tf.zeros(n), epochs=100, batch_size=1) y_NN = model.predict(x).flatten() plt.plot(x, y_NN, color='blue') plt.plot(x, y_e, color='red') plt.title('NN (blue) and exp (red)')
背景:希望用神经网络近似常微分方程(ODE)的解,上述是“零阶ODE”y(x)=exp(x)的极简示例。
错误分析与解决
核心错误原因
全局数组导致批次不匹配
你在loss函数中直接使用了全局numpy数组y_e,但训练采用batch_size=1的分批模式。每次训练只传入一个样本的x和y_pred,但loss计算时却是拿这个单样本的y_pred和整个数据集的y_e做差求平均,这完全不是单个样本的误差,而是全局平均误差。模型失去输入-输出映射的学习目标
这种全局误差的计算方式,会让模型趋向于输出一个固定常数——因为只有当y_pred等于y_e的全局平均值时,整体的平方误差才能最小,模型自然无法学习到x和exp(x)之间的映射关系。
修正方案
在loss函数中基于当前批次的输入x计算对应真实值,而非依赖全局数组。这样每个批次的误差计算都和当前样本的x绑定,模型才能学习到正确的映射:
import tensorflow as tf import numpy as np import matplotlib.pyplot as plt n = 101 x = np.linspace(start=0, stop=1, num=n) y_e = np.exp(x) model = tf.keras.models.Sequential([ tf.keras.layers.Dense(units=1, input_shape=[1]), tf.keras.layers.Dense(units=50, activation="softmax"), tf.keras.layers.Dense(units=1) ]) # 修正loss:基于当前批次的输入x计算对应的exp(x) def loss(y_true, y_pred): # 获取当前批次的输入张量x x_batch = model.inputs[0] # 计算当前批次每个样本的真实值exp(x) true_y = tf.exp(x_batch) # 计算当前批次的均方误差 L_ode = tf.reduce_mean(tf.square(y_pred - true_y), axis=-1) return L_ode model.compile('adam', loss) # 这里y_true可传任意值(代码中未使用),按需求传入tf.zeros(n) model.fit(x, tf.zeros(n), epochs=100, batch_size=1) y_NN = model.predict(x).flatten() plt.plot(x, y_NN, color='blue') plt.plot(x, y_e, color='red') plt.title('NN (blue) and exp (red)') plt.show()
补充说明
对于ODE近似场景,更规范的做法是让loss直接关联输入x和模型输出的微分关系(比如一阶ODE的y'=y),但核心逻辑一致:loss必须基于当前批次的输入计算对应约束,而非使用全局固定数据。
内容的提问来源于stack exchange,提问作者Marcel Steiner-Curtis
相关产品推荐
相关产品推荐

