You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 1.15神经网络损失值不下降问题求助

TensorFlow 1.x 二分类任务损失不下降问题排查与修复

我是TensorFlow新手,使用TensorFlow 1.1.15版本在乳腺癌数据集(含30个特征)上训练神经网络,网络架构为:
input_layer[30] --> hidden1[32] --> hidden2[16] --> output[1]
激活函数设置:

  • hidden1: relu
  • hidden2: relu
  • output: sigmoid

训练时使用tf.nn.softmax_cross_entropy_with_logits作为损失函数,但所有迭代中损失值完全保持不变,没有下降趋势。

原始代码

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
import tensorflow as tf

data = load_breast_cancer()
x_train, x_test, y_train, y_test = train_test_split(data.data, data.target, test_size=0.296, random_state=0)

n_input = 30
n_hide1 = 32
n_hide2 = 16
n_output = 1

weight_map = {
        "h1": tf.Variable(tf.random_normal(shape=(n_input, n_hide1))),
        "h2": tf.Variable(tf.random_normal(shape=(n_hide1, n_hide2))),
        "out": tf.Variable(tf.random_normal(shape=(n_hide2, n_output)))
}

bias_map = {
        "h1": tf.Variable(tf.random_normal(shape=(n_hide1, ))),
        "h2": tf.Variable(tf.random_normal(shape=(n_hide2, ))),
        "out": tf.Variable(tf.random_normal(shape=(n_output, ))),
}


def forward_propagation(x_inp, weights, biases):
    h1 = tf.nn.relu(tf.add(tf.matmul(x_inp, weights["h1"]), biases["h1"]))
    h2 = tf.nn.relu(tf.add(tf.matmul(h1, weights["h2"]), biases["h2"]))
    out = tf.nn.sigmoid(tf.add(tf.matmul(h2, weights["out"]), biases["out"]))
    return out


ALPHA = 0.01
x = tf.placeholder(dtype=tf.float32, shape=[None, 30])
y = tf.placeholder(dtype=tf.int32, shape=[None, ])

# STEP1: forward propagation
predictions = forward_propagation(x, weight_map, bias_map)

# STEP2: calculate cost func
cost = tf.nn.softmax_cross_entropy_with_logits(logits=predictions, labels=y)
cost_reduce_mean = tf.reduce_mean(cost)

# STEP3: backward propagation
optimizer = tf.train.AdamOptimizer(learning_rate=ALPHA)
optimize = optimizer.minimize(cost_reduce_mean)

with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    num_iter = 10
    for iteration in range(num_iter):
        c, _ = sess.run([cost_reduce_mean, optimize], feed_dict={x: x_train, y: y_train})
        print("ITER[{}]: {}".format(iteration, c))

输出情况:所有迭代中损失值完全没有变化。


问题根源分析

  1. 损失函数与任务不匹配
    你的任务是二分类,却使用了多分类专用的tf.nn.softmax_cross_entropy_with_logits,且该函数要求输入是未经过激活的原始输出(logits),而非sigmoid处理后的概率值,双重激活会彻底破坏梯度传递。

  2. 标签维度不兼容
    输出层维度是[None,1],但输入的标签是一维数组[None,],维度不匹配会导致损失计算异常,无法驱动权重更新。

  3. 数据未标准化
    乳腺癌数据集的特征尺度差异极大,未标准化的输入会导致神经网络收敛困难。

  4. 迭代次数与学习率设置不合理
    仅10次迭代不足以观察收敛趋势,0.01的学习率对Adam优化器来说过大,可能导致梯度爆炸。


修复后的代码

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
import tensorflow as tf

data = load_breast_cancer()
x_train, x_test, y_train, y_test = train_test_split(data.data, data.target, test_size=0.296, random_state=0)

# 数据标准化,消除特征尺度差异
scaler = StandardScaler()
x_train = scaler.fit_transform(x_train)
x_test = scaler.transform(x_test)

n_input = 30
n_hide1 = 32
n_hide2 = 16
n_output = 1

weight_map = {
        "h1": tf.Variable(tf.random_normal(shape=(n_input, n_hide1))),
        "h2": tf.Variable(tf.random_normal(shape=(n_hide1, n_hide2))),
        "out": tf.Variable(tf.random_normal(shape=(n_hide2, n_output)))
}

bias_map = {
        "h1": tf.Variable(tf.random_normal(shape=(n_hide1, ))),
        "h2": tf.Variable(tf.random_normal(shape=(n_hide2, ))),
        "out": tf.Variable(tf.random_normal(shape=(n_output, ))),
}


def forward_propagation(x_inp, weights, biases):
    h1 = tf.nn.relu(tf.add(tf.matmul(x_inp, weights["h1"]), biases["h1"]))
    h2 = tf.nn.relu(tf.add(tf.matmul(h1, weights["h2"]), biases["h2"]))
    # 返回未激活的logits,交给损失函数处理
    out_logits = tf.add(tf.matmul(h2, weights["out"]), biases["out"])
    return out_logits


ALPHA = 0.001  # 使用Adam默认学习率
x = tf.placeholder(dtype=tf.float32, shape=[None, 30])
# 标签转为float32并扩展维度,匹配输出层
y = tf.placeholder(dtype=tf.float32, shape=[None, 1])

# 前向传播
logits = forward_propagation(x, weight_map, bias_map)

# 替换为二分类专用损失函数
cost = tf.nn.sigmoid_cross_entropy_with_logits(logits=logits, labels=y)
cost_reduce_mean = tf.reduce_mean(cost)

# 反向传播优化
optimizer = tf.train.AdamOptimizer(learning_rate=ALPHA)
optimize = optimizer.minimize(cost_reduce_mean)

with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    num_iter = 100  # 增加迭代次数
    for iteration in range(num_iter):
        # 标签扩展维度后传入
        c, _ = sess.run([cost_reduce_mean, optimize], feed_dict={x: x_train, y: y_train.reshape(-1,1)})
        if (iteration+1) % 10 == 0:
            print("ITER[{}]: {}".format(iteration+1, c))

关键修改说明

  • 替换损失函数:用tf.nn.sigmoid_cross_entropy_with_logits适配二分类任务,该函数内部会处理sigmoid激活,避免手动激活导致的梯度问题。
  • 调整输出层:前向传播返回未激活的logits,保证梯度正常传递。
  • 标签维度匹配:将一维标签转为二维数组,与输出层维度一致。
  • 数据标准化:统一特征尺度,加速模型收敛。
  • 优化超参数:调整学习率为Adam默认的0.001,增加迭代次数到100次,便于观察损失下降趋势。

内容的提问来源于stack exchange,提问作者Ritik Kamra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 12:45:34