TensorFlow单隐藏层MNIST分类模型精度无法提升问题排查
问题诊断与修复方案
嘿,我一眼就看到你代码里的致命问题了——你踩了神经网络初始化的经典坑!
核心问题:全零权重初始化导致神经元“集体死亡”
你用tf.zeros初始化了所有权重w1和w2,这直接让整个网络彻底失效:
- 当
w1全为0时,隐藏层的输出z = X*w1 + b1结果全是0,经过ReLU激活后a = tf.nn.relu(z)也全是0(因为ReLU对0的输出还是0) - 接着后面的
logits = tf.matmul(a, w2) + b2自然也全是0,softmax之后每个类别的概率都是1/10=0.1,这就完美解释了为什么精度固定在~0.113,交叉熵损失固定在-ln(0.1)≈2.3
除此之外,还有两个小问题需要修正:
- 代码里缺少加载MNIST数据集的关键语句,程序会直接报错
- 测试阶段的
sess.run([correct_preds, accuracy])写法冗余,没必要同时运行两个张量,直接取accuracy即可
修正后的完整代码
import os import numpy as np import tensorflow as tf from tensorflow.examples.tutorials.mnist import input_data import time learning_rate = 0.01 batch_size = 128 n_epochs = 10 # 新增:加载MNIST数据集,one_hot=True对应你的Y占位符格式 mnist = input_data.read_data_sets('./data/mnist', one_hot=True) X = tf.placeholder(tf.float32, shape=(batch_size, 784)) Y = tf.placeholder(tf.float32, shape=(batch_size, 10)) # 改用随机正态分布初始化权重,打破对称性,避免神经元死亡 w1 = tf.Variable(tf.random_normal([X.shape[1], 30], stddev=0.01)) b1 = tf.Variable(tf.zeros([1, 30])) z = tf.matmul(X,w1) + b1 a = tf.nn.relu(z) w2 = tf.Variable(tf.random_normal([30, 10], stddev=0.01)) b2 = tf.Variable(tf.zeros([1, 10])) logits = tf.matmul(a,w2) + b2 entropy = tf.nn.softmax_cross_entropy_with_logits(logits = logits, labels = Y) loss = tf.reduce_mean(entropy) optimizer = tf.train.GradientDescentOptimizer(learning_rate).minimize(loss) with tf.Session() as sess: start_time = time.time() sess.run(tf.global_variables_initializer()) n_batches = int(mnist.train.num_examples/batch_size) for i in range(n_epochs): # train the model n_epochs times total_loss = 0 for _ in range(n_batches): X_batch, Y_batch = mnist.train.next_batch(batch_size) _, loss_batch = sess.run([optimizer, loss], feed_dict={X: X_batch, Y:Y_batch}) total_loss += loss_batch print('Average loss epoch {0}: {1}'.format(i, total_loss/n_batches)) print('Optimization Finished!') # should be around 0.35 after 25 epochs preds = tf.nn.softmax(logits) correct_preds = tf.equal(tf.argmax(preds, 1), tf.argmax(Y, 1)) accuracy = tf.reduce_sum(tf.cast(correct_preds, tf.float32)) n_batches = int(mnist.test.num_examples/batch_size) total_correct_preds = 0 for i in range(n_batches): X_batch, Y_batch = mnist.test.next_batch(batch_size) # 修正:直接运行accuracy即可,不需要同时跑correct_preds accuracy_batch = sess.run(accuracy, feed_dict={X: X_batch, Y:Y_batch}) total_correct_preds += accuracy_batch print('Accuracy {0}'.format(total_correct_preds/mnist.test.num_examples))
额外优化小建议
- 可以试试用
tf.contrib.layers.xavier_initializer()来初始化权重,这是专门针对ReLU激活的初始化方法,能让网络收敛更快 - 学习率可以适当调大(比如0.1),或者换成Adam优化器,训练效率会比梯度下降高不少
- 把训练轮数增加到25轮左右,精度应该能轻松达到95%以上
内容的提问来源于stack exchange,提问作者silent_dev
相关产品推荐
相关产品推荐

