You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow单隐藏层MNIST分类模型精度无法提升问题排查

问题诊断与修复方案

嘿,我一眼就看到你代码里的致命问题了——你踩了神经网络初始化的经典坑!

核心问题:全零权重初始化导致神经元“集体死亡”

你用tf.zeros初始化了所有权重w1和w2,这直接让整个网络彻底失效:

  • 当w1全为0时,隐藏层的输出z = X*w1 + b1结果全是0,经过ReLU激活后a = tf.nn.relu(z)也全是0(因为ReLU对0的输出还是0)
  • 接着后面的logits = tf.matmul(a, w2) + b2自然也全是0,softmax之后每个类别的概率都是1/10=0.1,这就完美解释了为什么精度固定在~0.113,交叉熵损失固定在-ln(0.1)≈2.3

除此之外,还有两个小问题需要修正:

  1. 代码里缺少加载MNIST数据集的关键语句,程序会直接报错
  2. 测试阶段的sess.run([correct_preds, accuracy])写法冗余,没必要同时运行两个张量,直接取accuracy即可

修正后的完整代码

import os
import numpy as np
import tensorflow as tf
from tensorflow.examples.tutorials.mnist import input_data
import time

learning_rate = 0.01
batch_size = 128
n_epochs = 10

# 新增:加载MNIST数据集,one_hot=True对应你的Y占位符格式
mnist = input_data.read_data_sets('./data/mnist', one_hot=True)

X = tf.placeholder(tf.float32, shape=(batch_size, 784))
Y = tf.placeholder(tf.float32, shape=(batch_size, 10))

# 改用随机正态分布初始化权重,打破对称性,避免神经元死亡
w1 = tf.Variable(tf.random_normal([X.shape[1], 30], stddev=0.01))
b1 = tf.Variable(tf.zeros([1, 30]))
z = tf.matmul(X,w1) + b1
a = tf.nn.relu(z)

w2 = tf.Variable(tf.random_normal([30, 10], stddev=0.01))
b2 = tf.Variable(tf.zeros([1, 10]))
logits = tf.matmul(a,w2) + b2

entropy = tf.nn.softmax_cross_entropy_with_logits(logits = logits, labels = Y)
loss = tf.reduce_mean(entropy)
optimizer = tf.train.GradientDescentOptimizer(learning_rate).minimize(loss)

with tf.Session() as sess:
    start_time = time.time()
    sess.run(tf.global_variables_initializer())
    n_batches = int(mnist.train.num_examples/batch_size)
    for i in range(n_epochs): # train the model n_epochs times
        total_loss = 0
        for _ in range(n_batches):
            X_batch, Y_batch = mnist.train.next_batch(batch_size)
            _, loss_batch = sess.run([optimizer, loss], feed_dict={X: X_batch, Y:Y_batch})
            total_loss += loss_batch
        print('Average loss epoch {0}: {1}'.format(i, total_loss/n_batches))
    print('Optimization Finished!') # should be around 0.35 after 25 epochs

    preds = tf.nn.softmax(logits)
    correct_preds = tf.equal(tf.argmax(preds, 1), tf.argmax(Y, 1))
    accuracy = tf.reduce_sum(tf.cast(correct_preds, tf.float32))

    n_batches = int(mnist.test.num_examples/batch_size)
    total_correct_preds = 0
    for i in range(n_batches):
        X_batch, Y_batch = mnist.test.next_batch(batch_size)
        # 修正:直接运行accuracy即可,不需要同时跑correct_preds
        accuracy_batch = sess.run(accuracy, feed_dict={X: X_batch, Y:Y_batch})
        total_correct_preds += accuracy_batch
    print('Accuracy {0}'.format(total_correct_preds/mnist.test.num_examples))

额外优化小建议

  • 可以试试用tf.contrib.layers.xavier_initializer()来初始化权重,这是专门针对ReLU激活的初始化方法,能让网络收敛更快
  • 学习率可以适当调大(比如0.1),或者换成Adam优化器,训练效率会比梯度下降高不少
  • 把训练轮数增加到25轮左右,精度应该能轻松达到95%以上

内容的提问来源于stack exchange,提问作者silent_dev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:55:04