神经网络拟合平方函数输出全为0问题排查与解决求助
问题描述
尝试构建神经网络拟合-50到50范围内数字的平方值,编写代码后所有输入的输出均为[[0.]],同时代码执行时出现以下提示:
oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable
TF_ENABLE_ONEDNN_OPTS=0.
原代码如下:
import tensorflow as tf import numpy as np from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense from tensorflow.keras import regularizers x_train = np.random.random((10000,1))*100-50 y_train = np.square(x_train) model = Sequential( [ Dense(8, activation = 'relu', kernel_regularizer = regularizers.l2(0.001), input_shape = (1,)), Dense(8, activation = 'relu', kernel_regularizer = regularizers.l2(0.001)), Dense(1, activation = 'relu') ] ) batch_size = 32 epochs = 100 model.compile(loss = 'mse', optimizer='adam') model.fit(x_train, y_train, batch_size=batch_size, epochs=epochs, verbose = 1) x = "n" while True: print("enter num:") x = input() if x == "end": break X = int(x) predicted_sum = model.predict(np.array([X])) print(predicted_sum)
原因分析与解决办法
输出层激活函数错误
最后一层使用relu激活函数是核心问题:ReLU的特性是当输入≤0时输出0。平方函数的输出虽为非负,但模型训练时因梯度不稳定等问题,输出层的原始预测值(激活前)可能为负,被ReLU截断为0。
解决:移除输出层的activation='relu',使用默认的线性激活(linear),回归任务需要连续输出值,无需非线性激活。数据未归一化导致训练不收敛
输入x的范围是[-50,50],输出y的范围是[0,2500],数值跨度大,会导致模型训练时梯度爆炸或消失,无法有效学习平方映射关系。
解决:对输入和输出做归一化处理,例如将x缩放到[-1,1],y缩放到[0,1],预测时再反归一化还原真实值。oneDNN提示与当前问题无关
该提示是TensorFlow启用oneDNN优化的通知,仅会导致微小的数值计算差异,和输出全为0的问题没有关联,无需处理。
修改后的示例代码
import tensorflow as tf import numpy as np from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense from tensorflow.keras import regularizers # 数据生成与归一化 x_train = np.random.random((10000,1))*100 - 50 y_train = np.square(x_train) # 归一化x到[-1,1],y到[0,1] x_mean = x_train.mean() x_std = x_train.std() x_train_normalized = (x_train - x_mean) / x_std y_max = y_train.max() y_train_normalized = y_train / y_max # 构建模型:移除输出层ReLU model = Sequential([ Dense(16, activation='relu', kernel_regularizer=regularizers.l2(0.001), input_shape=(1,)), Dense(16, activation='relu', kernel_regularizer=regularizers.l2(0.001)), Dense(1) # 默认线性激活 ]) batch_size = 32 epochs = 200 model.compile(loss='mse', optimizer='adam') model.fit(x_train_normalized, y_train_normalized, batch_size=batch_size, epochs=epochs, verbose=1) # 预测逻辑:加入反归一化 while True: print("enter num:") x_input = input() if x_input == "end": break X = int(x_input) # 对输入做归一化 X_normalized = (X - x_mean) / x_std predicted_normalized = model.predict(np.array([X_normalized]), verbose=0) # 反归一化得到真实平方值 predicted_value = predicted_normalized * y_max print(predicted_value)
内容的提问来源于stack exchange,提问作者harry

