TensorFlow MNIST专家示例准确率无法突破30%求助
Hey there, let's figure out why your MNIST model isn't hitting the expected accuracy. After going through your code, I've spotted several critical issues that are stopping your model from learning properly:
Key Problems Holding You Back
1. Wrong Learning Rate for Adadelta Optimizer
You’re using AdadeltaOptimizer with a learning rate of 0.0001, which is way too small for this optimizer. Adadelta is designed to auto-adjust its own learning rate, and its default value is 1.0. Using such a tiny learning rate means your model’s weights barely update during training—this is the primary reason your accuracy is stuck below 30%.
The official MNIST expert example uses AdamOptimizer(1e-4), which is a much better fit here. If you want to stick with Adadelta, set the learning rate to 1.0; otherwise, switching to Adam will give you more reliable results.
2. Dropout Is Disabled During Training
You’re passing keep_prob: 1 during training, which means the dropout layer does nothing (all neurons are kept). Dropout is meant to regularize the model during training by randomly disabling neurons—you should set keep_prob: 0.5 while training, and only use 1.0 when evaluating on the test set.
3. Unverified Convolution/Pooling Functions
You call conv_2d and max_pool_2x2 but haven’t included their code. If these functions are implemented incorrectly (e.g., wrong padding, stride, or kernel size), your model won’t process images properly. For reference, here are the official implementations you should match:
def conv_2d(x, W): return tf.nn.conv2d(x, W, strides=[1, 1, 1, 1], padding='SAME') def max_pool_2x2(x): return tf.nn.max_pool(x, ksize=[1, 2, 2, 1], strides=[1, 2, 2, 1], padding='SAME')
Using padding='VALID' or incorrect strides would break the feature extraction pipeline entirely.
Corrected Code Snippet
Here’s how to fix the main issues in your implementation:
def main(args): mnist = input_data.read_data_sets("/MNIST_data/", one_hot=True) x = tf.placeholder(tf.float32, shape=[None, 784]) y_ = tf.placeholder(tf.float32, shape=[None, 10]) x_image = tf.reshape(x, [-1, 28, 28, 1]) # Add missing conv/pool/weight/bias helper functions def conv_2d(x, W): return tf.nn.conv2d(x, W, strides=[1, 1, 1, 1], padding='SAME') def max_pool_2x2(x): return tf.nn.max_pool(x, ksize=[1, 2, 2, 1], strides=[1, 2, 2, 1], padding='SAME') def weight_variable(shape): initial = tf.truncated_normal(shape, stddev=0.1) return tf.Variable(initial) def bias_variable(shape): initial = tf.constant(0.1, shape=shape) return tf.Variable(initial) # Convolutional Layer 1 w_conv1 = weight_variable([5, 5, 1, 32]) b_conv1 = bias_variable([32]) h_conv1 = tf.nn.relu(conv_2d(x_image, w_conv1) + b_conv1) h_pool1 = max_pool_2x2(h_conv1) # Convolutional Layer 2 w_conv2 = weight_variable([5, 5, 32, 64]) b_conv2 = bias_variable([64]) h_conv2 = tf.nn.relu(conv_2d(h_pool1, w_conv2) + b_conv2) h_pool2 = max_pool_2x2(h_conv2) # Densely Connected Layer w_fc1 = weight_variable([7 * 7 * 64, 1024]) b_fc1 = bias_variable([1024]) h_pool2_flat = tf.reshape(h_pool2, [-1, 7 * 7 * 64]) h_fc1 = tf.nn.relu(tf.matmul(h_pool2_flat, w_fc1) + b_fc1) # Dropout Layer keep_prob = tf.placeholder(tf.float32) h_fc1_drop = tf.nn.dropout(h_fc1, keep_prob) # Output Layer w_fc2 = weight_variable([1024, 10]) b_fc2 = bias_variable([10]) y_conv = tf.matmul(h_fc1_drop, w_fc2) + b_fc2 # Loss and Optimizer (switched to Adam as in official example) cross_entropy = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(labels=y_, logits=y_conv)) train_step = tf.train.AdamOptimizer(1e-4).minimize(cross_entropy) # Accuracy Metrics correct_prediction = tf.equal(tf.argmax(y_conv, 1), tf.argmax(y_, 1)) accuracy = tf.reduce_mean(tf.cast(correct_prediction, tf.float32)) config = tf.ConfigProto(device_count={'GPU': 0}) with tf.Session(config=config) as sess: sess.run(tf.global_variables_initializer()) for i in range(20000): batch = mnist.train.next_batch(50) # Enable dropout during training train_step.run(feed_dict={x: batch[0], y_: batch[1], keep_prob: 0.5}) if i % 100 == 0: train_accuracy = accuracy.eval(feed_dict={ x: batch[0], y_: batch[1], keep_prob: 1.0 }) print("step %d, training accuracy %g" %(i, train_accuracy)) print("test accuracy %g" %accuracy.eval(feed_dict={ x: mnist.test.images, y_: mnist.test.labels, keep_prob: 1.0})) if __name__ == '__main__': tf.app.run(main = main)
Final Notes
After making these changes, your model should start learning correctly and reach the ~99% accuracy promised by the official example. The biggest fixes are switching to the right optimizer/learning rate and enabling dropout during training—those alone will make a massive difference.
内容的提问来源于stack exchange,提问作者rohanmishra

