You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow与sklearn的MLP模型性能差距疑问求助

Why Does sklearn's MLP Outperform TensorFlow's in My Experiments?

I recently ran a comparison between TensorFlow and sklearn's MLP implementations using two datasets: sklearn's built-in toy digits dataset and the MNIST dataset. Both models are single-hidden-layer MLPs with 300 neurons, and I tried to align all parameters as closely as possible (learning rate, L2 regularization, training epochs, batch size, Adam optimizer settings). However, sklearn's model consistently achieved higher accuracy:

  • On the sklearn digits dataset: ~83% accuracy with TensorFlow vs ~90% with sklearn
  • On MNIST: ~94% accuracy with TensorFlow vs ~97% with sklearn

I'm wondering if sklearn has special optimizations for its MLP that I'm missing in my TensorFlow code. Here's the full code I used:

from __future__ import print_function
# Import MNIST data
from tensorflow.examples.tutorials.mnist import input_data
mnist = input_data.read_data_sets("/tmp/data/", one_hot=True)
import tensorflow as tf
from sklearn.datasets import load_digits
import numpy as np

x_train = mnist.train.images
y_train = mnist.train.labels
x_test = mnist.test.images
y_test = mnist.test.labels
n_train = mnist.train.images.shape[0]

# Parameters
learning_rate = 1e-3
lambda_val = 1e-5
training_epochs = 30
batch_size = 200
display_step = 1
# Network Parameters
n_hidden_1 = 300 # 1st layer number of neurons
n_input = x_train.shape[1] # MNIST data input (img shape: 28*28)
n_classes = 10 # MNIST total classes (0-9 digits)

# tf Graph input
X = tf.placeholder("float", [None, n_input])
Y = tf.placeholder("float", [None, n_classes])

# Store layers weight & bias
weights = {
    'h1': tf.Variable(tf.random_normal([n_input, n_hidden_1])),
    'out': tf.Variable(tf.random_normal([n_hidden_1, n_classes]))
}
biases = {
    'b1': tf.Variable(tf.random_normal([n_hidden_1])),
    'out': tf.Variable(tf.random_normal([n_classes]))
}

# Create model
def multilayer_perceptron(x):
    # Hidden fully connected layer with 256 neurons
    layer_1 = tf.add(tf.matmul(x, weights['h1']), biases['b1'])
    # Activation
    layer_1_relu = tf.nn.relu(layer_1)
    # Output fully connected layer with a neuron for each class
    out_layer = tf.matmul(layer_1_relu, weights['out']) + biases['out']
    return out_layer

# Construct model
logits = multilayer_perceptron(X)

# Define loss and optimizer
loss_op = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=logits, labels=Y)) + lambda_val*tf.nn.l2_loss(weights['h1']) + lambda_val*tf.nn.l2_loss(weights['out'])
optimizer = tf.train.AdamOptimizer(learning_rate=learning_rate)
train_op = optimizer.minimize(loss_op)

# Test model
pred = tf.nn.softmax(logits) # Apply softmax to logits
correct_prediction = tf.equal(tf.argmax(pred, 1), tf.argmax(Y, 1))
# Calculate accuracy
accuracy = tf.reduce_mean(tf.cast(correct_prediction, "float"))

# Initializing the variables
init = tf.global_variables_initializer()

with tf.Session() as sess:
    sess.run(init)
    # Training cycle
    for epoch in range(training_epochs):
        avg_cost = 0.
        total_batch = int(n_train/batch_size)
        # Loop over all batches
        ptr = 0
        for i in range(total_batch):
            next_ptr = ptr + batch_size
            if next_ptr > len(x_train):
                next_ptr = len(x_train)
            batch_x, batch_y = x_train[ptr:next_ptr],y_train[ptr:next_ptr]
            ptr += batch_size
            # Run optimization op (backprop) and cost op (to get loss value)
            _, c = sess.run([train_op, loss_op], feed_dict={X: batch_x, Y: batch_y})
            # Compute average loss
            avg_cost += c / total_batch
        # Display logs per epoch step
        if epoch % display_step == 0:
            print("Epoch:", '%04d' % (epoch+1), "cost={:.9f}".format(avg_cost))
    print("Optimization Finished!")
    print("Accuracy on training set: ", accuracy.eval({X:x_train,Y:y_train}))
    print("Accuracy on testing set:", accuracy.eval({X: x_test, Y: y_test}))

print("Experimenting sklearn...")
# now experiment with sklearn
from sklearn.datasets import load_digits
import numpy as np
from sklearn.neural_network import MLPClassifier
import time

# use MLP
t_start = time.time()
print('fitting MLP...')
clf = MLPClassifier(solver='adam', alpha=1e-5, hidden_layer_sizes=(300,),max_iter=training_epochs)
clf.fit(x_train,y_train)
print('fitted MLP in {:.2f} seconds'.format(time.time() - t_start))
print('predicting...')
labels_predicted = clf.predict(x_test)
print('accuracy: {:.2f} %'.format(np.mean(np.argmax(y_test,axis=1) == np.argmax(labels_predicted,axis=1)) * 100))

Key Differences Explaining the Performance Gap

Let's break down the likely reasons behind the discrepancy, starting with the most critical issues in your code:

  1. Incorrect Label Format for sklearn MLP
    This is a major mistake in your code: sklearn's MLPClassifier expects integer-form class labels (shape (n_samples,), e.g., [0,1,2,...,9]) instead of one-hot encoded labels (shape (n_samples, 10)). When you pass one-hot labels to clf.fit(), sklearn treats this as a multi-label classification task (predicting multiple classes per sample) instead of the multi-class task you intended. Your accuracy calculation using np.argmax() on the sklearn model's output is not appropriate for multi-label classification, which means the reported sklearn accuracy is either incorrect or the model is behaving in an unintended way.

    Fix this by converting your one-hot labels back to integer labels before passing to sklearn:

    y_train_sklearn = np.argmax(y_train, axis=1)
    y_test_sklearn = np.argmax(y_test, axis=1)
    clf.fit(x_train, y_train_sklearn)
    # Then calculate accuracy correctly:
    labels_predicted = clf.predict(x_test)
    print('accuracy: {:.2f} %'.format(np.mean(y_test_sklearn == labels_predicted) * 100))
    
  2. Weight Initialization Differences
    Your TensorFlow code uses tf.random_normal() (default mean=0, stddev=1) to initialize weights, which can lead to excessively large initial values. For ReLU-activated layers, this often causes "dead neurons" (neurons that never activate because their inputs are always negative), slowing down training and hurting performance.

    Sklearn's MLPClassifier uses Glorot uniform initialization by default (scaling weights based on the number of input and output neurons), which is far more stable for deep learning models. To match this in TensorFlow, replace your weight initialization code with:

    weights = {
        'h1': tf.Variable(tf.glorot_uniform_initializer()([n_input, n_hidden_1])),
        'out': tf.Variable(tf.glorot_uniform_initializer()([n_hidden_1, n_classes]))
    }
    
  3. Training Data Shuffling
    Your TensorFlow code processes training batches in sequential order without shuffling the data each epoch. This can lead to the model learning spurious patterns from the order of the data, reducing generalization.

    Sklearn's MLPClassifier defaults to shuffle=True, which shuffles the training data at the start of each epoch. This ensures the model sees varied batches every epoch, leading to better convergence and higher accuracy. To add this to your TensorFlow code, shuffle your training data each epoch:

    for epoch in range(training_epochs):
        # Shuffle training data each epoch
        permutation = np.random.permutation(n_train)
        x_train_shuffled = x_train[permutation]
        y_train_shuffled = y_train[permutation]
        avg_cost = 0.
        total_batch = int(n_train/batch_size)
        ptr = 0
        for i in range(total_batch):
            next_ptr = ptr + batch_size
            if next_ptr > len(x_train_shuffled):
                next_ptr = len(x_train_shuffled)
            batch_x, batch_y = x_train_shuffled[ptr:next_ptr], y_train_shuffled[ptr:next_ptr]
            ptr += batch_size
            _, c = sess.run([train_op, loss_op], feed_dict={X: batch_x, Y: batch_y})
            avg_cost += c / total_batch
    
  4. Minor Implementation Nuances
    Even with all parameters aligned, there can be subtle differences in how each library implements the Adam optimizer, loss function averaging, and regularization. For example:

    • Sklearn may use slightly different gradient clipping or numerical stability checks under the hood.
    • TensorFlow's tf.nn.softmax_cross_entropy_with_logits() computes the loss per sample and averages it, while sklearn's loss calculation is mathematically equivalent but may use different floating-point precision handling.

After Fixing These Issues...

Once you correct the label format, weight initialization, and add data shuffling to your TensorFlow code, you should see the performance gap narrow significantly. Sklearn's MLP is well-optimized with sensible defaults, but TensorFlow can match its performance when configured properly.

内容的提问来源于stack exchange,提问作者lenhhoxung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:02:23