TensorFlow 1.15神经网络损失值不下降问题求助
TensorFlow 1.x 二分类任务损失不下降问题排查与修复
我是TensorFlow新手,使用TensorFlow 1.1.15版本在乳腺癌数据集(含30个特征)上训练神经网络,网络架构为:input_layer[30] --> hidden1[32] --> hidden2[16] --> output[1]
激活函数设置:
- hidden1: relu
- hidden2: relu
- output: sigmoid
训练时使用tf.nn.softmax_cross_entropy_with_logits作为损失函数,但所有迭代中损失值完全保持不变,没有下降趋势。
原始代码
from sklearn.datasets import load_breast_cancer from sklearn.model_selection import train_test_split import tensorflow as tf data = load_breast_cancer() x_train, x_test, y_train, y_test = train_test_split(data.data, data.target, test_size=0.296, random_state=0) n_input = 30 n_hide1 = 32 n_hide2 = 16 n_output = 1 weight_map = { "h1": tf.Variable(tf.random_normal(shape=(n_input, n_hide1))), "h2": tf.Variable(tf.random_normal(shape=(n_hide1, n_hide2))), "out": tf.Variable(tf.random_normal(shape=(n_hide2, n_output))) } bias_map = { "h1": tf.Variable(tf.random_normal(shape=(n_hide1, ))), "h2": tf.Variable(tf.random_normal(shape=(n_hide2, ))), "out": tf.Variable(tf.random_normal(shape=(n_output, ))), } def forward_propagation(x_inp, weights, biases): h1 = tf.nn.relu(tf.add(tf.matmul(x_inp, weights["h1"]), biases["h1"])) h2 = tf.nn.relu(tf.add(tf.matmul(h1, weights["h2"]), biases["h2"])) out = tf.nn.sigmoid(tf.add(tf.matmul(h2, weights["out"]), biases["out"])) return out ALPHA = 0.01 x = tf.placeholder(dtype=tf.float32, shape=[None, 30]) y = tf.placeholder(dtype=tf.int32, shape=[None, ]) # STEP1: forward propagation predictions = forward_propagation(x, weight_map, bias_map) # STEP2: calculate cost func cost = tf.nn.softmax_cross_entropy_with_logits(logits=predictions, labels=y) cost_reduce_mean = tf.reduce_mean(cost) # STEP3: backward propagation optimizer = tf.train.AdamOptimizer(learning_rate=ALPHA) optimize = optimizer.minimize(cost_reduce_mean) with tf.Session() as sess: sess.run(tf.global_variables_initializer()) num_iter = 10 for iteration in range(num_iter): c, _ = sess.run([cost_reduce_mean, optimize], feed_dict={x: x_train, y: y_train}) print("ITER[{}]: {}".format(iteration, c))
输出情况:所有迭代中损失值完全没有变化。
问题根源分析
损失函数与任务不匹配
你的任务是二分类,却使用了多分类专用的tf.nn.softmax_cross_entropy_with_logits,且该函数要求输入是未经过激活的原始输出(logits),而非sigmoid处理后的概率值,双重激活会彻底破坏梯度传递。标签维度不兼容
输出层维度是[None,1],但输入的标签是一维数组[None,],维度不匹配会导致损失计算异常,无法驱动权重更新。数据未标准化
乳腺癌数据集的特征尺度差异极大,未标准化的输入会导致神经网络收敛困难。迭代次数与学习率设置不合理
仅10次迭代不足以观察收敛趋势,0.01的学习率对Adam优化器来说过大,可能导致梯度爆炸。
修复后的代码
from sklearn.datasets import load_breast_cancer from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler import tensorflow as tf data = load_breast_cancer() x_train, x_test, y_train, y_test = train_test_split(data.data, data.target, test_size=0.296, random_state=0) # 数据标准化,消除特征尺度差异 scaler = StandardScaler() x_train = scaler.fit_transform(x_train) x_test = scaler.transform(x_test) n_input = 30 n_hide1 = 32 n_hide2 = 16 n_output = 1 weight_map = { "h1": tf.Variable(tf.random_normal(shape=(n_input, n_hide1))), "h2": tf.Variable(tf.random_normal(shape=(n_hide1, n_hide2))), "out": tf.Variable(tf.random_normal(shape=(n_hide2, n_output))) } bias_map = { "h1": tf.Variable(tf.random_normal(shape=(n_hide1, ))), "h2": tf.Variable(tf.random_normal(shape=(n_hide2, ))), "out": tf.Variable(tf.random_normal(shape=(n_output, ))), } def forward_propagation(x_inp, weights, biases): h1 = tf.nn.relu(tf.add(tf.matmul(x_inp, weights["h1"]), biases["h1"])) h2 = tf.nn.relu(tf.add(tf.matmul(h1, weights["h2"]), biases["h2"])) # 返回未激活的logits,交给损失函数处理 out_logits = tf.add(tf.matmul(h2, weights["out"]), biases["out"]) return out_logits ALPHA = 0.001 # 使用Adam默认学习率 x = tf.placeholder(dtype=tf.float32, shape=[None, 30]) # 标签转为float32并扩展维度,匹配输出层 y = tf.placeholder(dtype=tf.float32, shape=[None, 1]) # 前向传播 logits = forward_propagation(x, weight_map, bias_map) # 替换为二分类专用损失函数 cost = tf.nn.sigmoid_cross_entropy_with_logits(logits=logits, labels=y) cost_reduce_mean = tf.reduce_mean(cost) # 反向传播优化 optimizer = tf.train.AdamOptimizer(learning_rate=ALPHA) optimize = optimizer.minimize(cost_reduce_mean) with tf.Session() as sess: sess.run(tf.global_variables_initializer()) num_iter = 100 # 增加迭代次数 for iteration in range(num_iter): # 标签扩展维度后传入 c, _ = sess.run([cost_reduce_mean, optimize], feed_dict={x: x_train, y: y_train.reshape(-1,1)}) if (iteration+1) % 10 == 0: print("ITER[{}]: {}".format(iteration+1, c))
关键修改说明
- 替换损失函数:用
tf.nn.sigmoid_cross_entropy_with_logits适配二分类任务,该函数内部会处理sigmoid激活,避免手动激活导致的梯度问题。 - 调整输出层:前向传播返回未激活的logits,保证梯度正常传递。
- 标签维度匹配:将一维标签转为二维数组,与输出层维度一致。
- 数据标准化:统一特征尺度,加速模型收敛。
- 优化超参数:调整学习率为Adam默认的0.001,增加迭代次数到100次,便于观察损失下降趋势。
内容的提问来源于stack exchange,提问作者Ritik Kamra
相关产品推荐
相关产品推荐

