TensorFlow与sklearn的MLP模型性能差距疑问求助
I recently ran a comparison between TensorFlow and sklearn's MLP implementations using two datasets: sklearn's built-in toy digits dataset and the MNIST dataset. Both models are single-hidden-layer MLPs with 300 neurons, and I tried to align all parameters as closely as possible (learning rate, L2 regularization, training epochs, batch size, Adam optimizer settings). However, sklearn's model consistently achieved higher accuracy:
- On the sklearn digits dataset: ~83% accuracy with TensorFlow vs ~90% with sklearn
- On MNIST: ~94% accuracy with TensorFlow vs ~97% with sklearn
I'm wondering if sklearn has special optimizations for its MLP that I'm missing in my TensorFlow code. Here's the full code I used:
from __future__ import print_function # Import MNIST data from tensorflow.examples.tutorials.mnist import input_data mnist = input_data.read_data_sets("/tmp/data/", one_hot=True) import tensorflow as tf from sklearn.datasets import load_digits import numpy as np x_train = mnist.train.images y_train = mnist.train.labels x_test = mnist.test.images y_test = mnist.test.labels n_train = mnist.train.images.shape[0] # Parameters learning_rate = 1e-3 lambda_val = 1e-5 training_epochs = 30 batch_size = 200 display_step = 1 # Network Parameters n_hidden_1 = 300 # 1st layer number of neurons n_input = x_train.shape[1] # MNIST data input (img shape: 28*28) n_classes = 10 # MNIST total classes (0-9 digits) # tf Graph input X = tf.placeholder("float", [None, n_input]) Y = tf.placeholder("float", [None, n_classes]) # Store layers weight & bias weights = { 'h1': tf.Variable(tf.random_normal([n_input, n_hidden_1])), 'out': tf.Variable(tf.random_normal([n_hidden_1, n_classes])) } biases = { 'b1': tf.Variable(tf.random_normal([n_hidden_1])), 'out': tf.Variable(tf.random_normal([n_classes])) } # Create model def multilayer_perceptron(x): # Hidden fully connected layer with 256 neurons layer_1 = tf.add(tf.matmul(x, weights['h1']), biases['b1']) # Activation layer_1_relu = tf.nn.relu(layer_1) # Output fully connected layer with a neuron for each class out_layer = tf.matmul(layer_1_relu, weights['out']) + biases['out'] return out_layer # Construct model logits = multilayer_perceptron(X) # Define loss and optimizer loss_op = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=logits, labels=Y)) + lambda_val*tf.nn.l2_loss(weights['h1']) + lambda_val*tf.nn.l2_loss(weights['out']) optimizer = tf.train.AdamOptimizer(learning_rate=learning_rate) train_op = optimizer.minimize(loss_op) # Test model pred = tf.nn.softmax(logits) # Apply softmax to logits correct_prediction = tf.equal(tf.argmax(pred, 1), tf.argmax(Y, 1)) # Calculate accuracy accuracy = tf.reduce_mean(tf.cast(correct_prediction, "float")) # Initializing the variables init = tf.global_variables_initializer() with tf.Session() as sess: sess.run(init) # Training cycle for epoch in range(training_epochs): avg_cost = 0. total_batch = int(n_train/batch_size) # Loop over all batches ptr = 0 for i in range(total_batch): next_ptr = ptr + batch_size if next_ptr > len(x_train): next_ptr = len(x_train) batch_x, batch_y = x_train[ptr:next_ptr],y_train[ptr:next_ptr] ptr += batch_size # Run optimization op (backprop) and cost op (to get loss value) _, c = sess.run([train_op, loss_op], feed_dict={X: batch_x, Y: batch_y}) # Compute average loss avg_cost += c / total_batch # Display logs per epoch step if epoch % display_step == 0: print("Epoch:", '%04d' % (epoch+1), "cost={:.9f}".format(avg_cost)) print("Optimization Finished!") print("Accuracy on training set: ", accuracy.eval({X:x_train,Y:y_train})) print("Accuracy on testing set:", accuracy.eval({X: x_test, Y: y_test})) print("Experimenting sklearn...") # now experiment with sklearn from sklearn.datasets import load_digits import numpy as np from sklearn.neural_network import MLPClassifier import time # use MLP t_start = time.time() print('fitting MLP...') clf = MLPClassifier(solver='adam', alpha=1e-5, hidden_layer_sizes=(300,),max_iter=training_epochs) clf.fit(x_train,y_train) print('fitted MLP in {:.2f} seconds'.format(time.time() - t_start)) print('predicting...') labels_predicted = clf.predict(x_test) print('accuracy: {:.2f} %'.format(np.mean(np.argmax(y_test,axis=1) == np.argmax(labels_predicted,axis=1)) * 100))
Key Differences Explaining the Performance Gap
Let's break down the likely reasons behind the discrepancy, starting with the most critical issues in your code:
Incorrect Label Format for sklearn MLP
This is a major mistake in your code: sklearn'sMLPClassifierexpects integer-form class labels (shape(n_samples,), e.g.,[0,1,2,...,9]) instead of one-hot encoded labels (shape(n_samples, 10)). When you pass one-hot labels toclf.fit(), sklearn treats this as a multi-label classification task (predicting multiple classes per sample) instead of the multi-class task you intended. Your accuracy calculation usingnp.argmax()on the sklearn model's output is not appropriate for multi-label classification, which means the reported sklearn accuracy is either incorrect or the model is behaving in an unintended way.Fix this by converting your one-hot labels back to integer labels before passing to sklearn:
y_train_sklearn = np.argmax(y_train, axis=1) y_test_sklearn = np.argmax(y_test, axis=1) clf.fit(x_train, y_train_sklearn) # Then calculate accuracy correctly: labels_predicted = clf.predict(x_test) print('accuracy: {:.2f} %'.format(np.mean(y_test_sklearn == labels_predicted) * 100))Weight Initialization Differences
Your TensorFlow code usestf.random_normal()(default mean=0, stddev=1) to initialize weights, which can lead to excessively large initial values. For ReLU-activated layers, this often causes "dead neurons" (neurons that never activate because their inputs are always negative), slowing down training and hurting performance.Sklearn's
MLPClassifieruses Glorot uniform initialization by default (scaling weights based on the number of input and output neurons), which is far more stable for deep learning models. To match this in TensorFlow, replace your weight initialization code with:weights = { 'h1': tf.Variable(tf.glorot_uniform_initializer()([n_input, n_hidden_1])), 'out': tf.Variable(tf.glorot_uniform_initializer()([n_hidden_1, n_classes])) }Training Data Shuffling
Your TensorFlow code processes training batches in sequential order without shuffling the data each epoch. This can lead to the model learning spurious patterns from the order of the data, reducing generalization.Sklearn's
MLPClassifierdefaults toshuffle=True, which shuffles the training data at the start of each epoch. This ensures the model sees varied batches every epoch, leading to better convergence and higher accuracy. To add this to your TensorFlow code, shuffle your training data each epoch:for epoch in range(training_epochs): # Shuffle training data each epoch permutation = np.random.permutation(n_train) x_train_shuffled = x_train[permutation] y_train_shuffled = y_train[permutation] avg_cost = 0. total_batch = int(n_train/batch_size) ptr = 0 for i in range(total_batch): next_ptr = ptr + batch_size if next_ptr > len(x_train_shuffled): next_ptr = len(x_train_shuffled) batch_x, batch_y = x_train_shuffled[ptr:next_ptr], y_train_shuffled[ptr:next_ptr] ptr += batch_size _, c = sess.run([train_op, loss_op], feed_dict={X: batch_x, Y: batch_y}) avg_cost += c / total_batchMinor Implementation Nuances
Even with all parameters aligned, there can be subtle differences in how each library implements the Adam optimizer, loss function averaging, and regularization. For example:- Sklearn may use slightly different gradient clipping or numerical stability checks under the hood.
- TensorFlow's
tf.nn.softmax_cross_entropy_with_logits()computes the loss per sample and averages it, while sklearn's loss calculation is mathematically equivalent but may use different floating-point precision handling.
After Fixing These Issues...
Once you correct the label format, weight initialization, and add data shuffling to your TensorFlow code, you should see the performance gap narrow significantly. Sklearn's MLP is well-optimized with sensible defaults, but TensorFlow can match its performance when configured properly.
内容的提问来源于stack exchange,提问作者lenhhoxung

