You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow自定义loss function问题:拟合指数函数为何收敛至常数?

问题描述

第一段代码可在[0,1]区间用神经网络近似指数函数exp(x),运行后能正常收敛:

#   code works
import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt
#   fit an exponential
n = 101
x = np.linspace(start=0, stop=1, num=n)
y_e = np.exp(x)
#   any odd neural net with sufficient degrees of freedom
model = tf.keras.models.Sequential([
     tf.keras.layers.Dense(units=1, input_shape=[1]),
     tf.keras.layers.Dense(units=50, activation="softmax"),
     tf.keras.layers.Dense(units=1)
])
#   loss function
def loss(y_true, y_pred): 
    L_ode = tf.reduce_mean(tf.square(y_pred - y_true), axis=-1)
    return L_ode
model.compile('adam', loss)
model.fit(x, y_e, epochs=100, batch_size=1)
y_NN = model.predict(x).flatten()
plt.plot(x, y_NN, color='blue')
plt.plot(x, y_e, color='red')
plt.title('NN (blue) and exp (red)')

修改loss函数并在model.fit中传入tf.zeros(n)作为y_true后,代码运行正常但模型收敛至常数,无法近似exp(x):

import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt
#   fit an exponential
n = 101
x = np.linspace(start=0, stop=1, num=n)
y_e = np.exp(x)
#   any odd neural net with sufficient degrees of freedom
model = tf.keras.models.Sequential([
     tf.keras.layers.Dense(units=1, input_shape=[1]),
     tf.keras.layers.Dense(units=50, activation="softmax"),
     tf.keras.layers.Dense(units=1)
])
#   loss function
def loss(y_true, y_pred): 
    L_ode = tf.reduce_mean(tf.square((y_pred - y_e) - y_true), axis=-1)
    return L_ode
model.compile('adam', loss)
model.fit(x, tf.zeros(n), epochs=100, batch_size=1)
y_NN = model.predict(x).flatten()
plt.plot(x, y_NN, color='blue')
plt.plot(x, y_e, color='red')
plt.title('NN (blue) and exp (red)')

背景:希望用神经网络近似常微分方程(ODE)的解,上述是“零阶ODE”y(x)=exp(x)的极简示例。

错误分析与解决

核心错误原因

  1. 全局数组导致批次不匹配
    你在loss函数中直接使用了全局numpy数组y_e,但训练采用batch_size=1的分批模式。每次训练只传入一个样本的x和y_pred,但loss计算时却是拿这个单样本的y_pred和整个数据集的y_e做差求平均,这完全不是单个样本的误差,而是全局平均误差。

  2. 模型失去输入-输出映射的学习目标
    这种全局误差的计算方式,会让模型趋向于输出一个固定常数——因为只有当y_pred等于y_e的全局平均值时,整体的平方误差才能最小,模型自然无法学习到x和exp(x)之间的映射关系。

修正方案

在loss函数中基于当前批次的输入x计算对应真实值,而非依赖全局数组。这样每个批次的误差计算都和当前样本的x绑定,模型才能学习到正确的映射:

import tensorflow as tf
import numpy as np
import matplotlib.pyplot as plt

n = 101
x = np.linspace(start=0, stop=1, num=n)
y_e = np.exp(x)

model = tf.keras.models.Sequential([
     tf.keras.layers.Dense(units=1, input_shape=[1]),
     tf.keras.layers.Dense(units=50, activation="softmax"),
     tf.keras.layers.Dense(units=1)
])

# 修正loss:基于当前批次的输入x计算对应的exp(x)
def loss(y_true, y_pred):
    # 获取当前批次的输入张量x
    x_batch = model.inputs[0]
    # 计算当前批次每个样本的真实值exp(x)
    true_y = tf.exp(x_batch)
    # 计算当前批次的均方误差
    L_ode = tf.reduce_mean(tf.square(y_pred - true_y), axis=-1)
    return L_ode

model.compile('adam', loss)
# 这里y_true可传任意值(代码中未使用),按需求传入tf.zeros(n)
model.fit(x, tf.zeros(n), epochs=100, batch_size=1)

y_NN = model.predict(x).flatten()
plt.plot(x, y_NN, color='blue')
plt.plot(x, y_e, color='red')
plt.title('NN (blue) and exp (red)')
plt.show()

补充说明

对于ODE近似场景,更规范的做法是让loss直接关联输入x和模型输出的微分关系(比如一阶ODE的y'=y),但核心逻辑一致:loss必须基于当前批次的输入计算对应约束,而非使用全局固定数据。

内容的提问来源于stack exchange,提问作者Marcel Steiner-Curtis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 14:18:18