TensorFlow新手求助:两层神经网络训练准确率无提升问题
Hey there! Let's work through why your TensorFlow two-layer network isn't improving accuracy with the UCI Heart Disease dataset. I’ll break down common pitfalls and fixes based on the code snippet you shared.
First: Fix the Incomplete Forward Pass
Your forwardprop function cuts off at h_2...—we need to finish defining the full network. Since you have three weight matrices (w_1, w_2, w_3), I assume you’re building an input → hidden layer 1 → hidden layer 2 → output structure. Here’s how to complete it:
def forwardprop(X, w_1, w_2, w_3): h_1 = tf.nn.sigmoid(tf.matmul(X, w_1)) h_2 = tf.nn.sigmoid(tf.matmul(h_1, w_2)) # Complete the second hidden layer y_pred = tf.nn.sigmoid(tf.matmul(h_2, w_3)) # Output layer for binary classification return y_pred
1. Standardize Your Data (Critical for Neural Networks)
The UCI Heart Disease dataset mixes numerical features (like age, cholesterol) and categorical ones (like sex, chest pain type). Neural networks (especially with sigmoid activation) are sensitive to unnormalized input ranges—values that are too large/small can cause gradient vanishing and slow learning.
Add this preprocessing step before splitting your data:
# Load and prepare data df = pd.read_csv('heart.csv') # Update with your dataset path X = df.drop('target', axis=1).values y = df['target'].values.reshape(-1, 1) # Reshape labels to match output layer dimensions # Standardize features to mean=0, std=1 from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # Split into train/test sets with scaled data X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2, random_state=RANDOM_SEED)
2. Improve Weight Initialization
Your current init_weights uses tf.random_normal(stddev=0.1), which isn’t ideal for sigmoid activation—this can lead to saturated neurons and vanishing gradients. Switch to Xavier initialization, which is designed for sigmoid/tanh layers:
def init_weights(shape): return tf.get_variable("weights", shape, initializer=tf.contrib.layers.xavier_initializer())
3. Swap Sigmoid for a Better Activation Function
Sigmoid is prone to gradient vanishing in even shallow networks. Replace it with ReLU (or LeakyReLU for edge cases) to keep gradients flowing:
def forwardprop(X, w_1, w_2, w_3): h_1 = tf.nn.relu(tf.matmul(X, w_1)) h_2 = tf.nn.relu(tf.matmul(h_1, w_2)) y_pred = tf.nn.sigmoid(tf.matmul(h_2, w_3)) # Keep sigmoid for binary output return y_pred
4. Use the Right Loss Function & Optimizer
For binary classification (heart disease presence/absence), avoid mean squared error. Use binary cross-entropy—and for better numerical stability, pass the raw logits (pre-sigmoid output) to the loss function:
# Define placeholders input_dim = X_train.shape[1] # 13 features for UCI Heart Disease hidden_dim1 = 64 # Adjust based on your needs hidden_dim2 = 32 output_dim = 1 X = tf.placeholder(tf.float32, shape=[None, input_dim]) y = tf.placeholder(tf.float32, shape=[None, output_dim]) # Initialize weights w_1 = init_weights((input_dim, hidden_dim1)) w_2 = init_weights((hidden_dim1, hidden_dim2)) w_3 = init_weights((hidden_dim2, output_dim)) # Forward pass (get logits for loss calculation) h_1 = tf.nn.relu(tf.matmul(X, w_1)) h_2 = tf.nn.relu(tf.matmul(h_1, w_2)) logits = tf.matmul(h_2, w_3) y_pred = tf.nn.sigmoid(logits) # Loss function: Binary cross-entropy with logits loss = tf.reduce_mean(tf.nn.sigmoid_cross_entropy_with_logits(labels=y, logits=logits)) # Use Adam optimizer instead of vanilla SGD (adapts learning rate automatically) optimizer = tf.train.AdamOptimizer(learning_rate=0.001).minimize(loss) # Accuracy metric correct_pred = tf.equal(tf.round(y_pred), y) accuracy = tf.reduce_mean(tf.cast(correct_pred, tf.float32))
5. Tune Training Parameters
- Epochs: If you’re only training for a few dozen epochs, your model might not have enough time to learn. Try 500–1000 epochs.
- Batch Size: Instead of training on the full dataset every time, use mini-batches (e.g., 32 or 64 samples) to stabilize learning:
epochs = 1000 batch_size = 32 with tf.Session() as sess: sess.run(tf.global_variables_initializer()) for epoch in range(epochs): # Train in batches for i in range(0, len(X_train), batch_size): batch_X = X_train[i:i+batch_size] batch_y = y_train[i:i+batch_size] sess.run(optimizer, feed_dict={X: batch_X, y: batch_y}) # Print progress every 10 epochs if epoch % 10 == 0: train_acc = sess.run(accuracy, feed_dict={X: X_train, y: y_train}) test_acc = sess.run(accuracy, feed_dict={X: X_test, y: y_test}) print(f"Epoch {epoch}: Train Acc = {train_acc:.4f}, Test Acc = {test_acc:.4f}")
6. Check for Class Imbalance
The UCI Heart Disease dataset has a roughly 50/50 split, but if your version is imbalanced, your model might favor the majority class. Fix this with weighted loss:
# Calculate class weights pos_count = np.sum(y_train == 1) neg_count = len(y_train) - pos_count pos_weight = neg_count / pos_count # Weight for positive class # Use weighted cross-entropy loss = tf.reduce_mean(tf.nn.weighted_cross_entropy_with_logits(labels=y, logits=logits, pos_weight=pos_weight))
Final Notes
- Adjust hidden layer sizes (try 32, 64, or 128) to find a balance between underfitting and overfitting.
- Add dropout if you see overfitting (train accuracy high, test accuracy low):
h_1 = tf.nn.dropout(tf.nn.relu(tf.matmul(X, w_1)), keep_prob=0.8) h_2 = tf.nn.dropout(tf.nn.relu(tf.matmul(h_1, w_2)), keep_prob=0.8)
内容的提问来源于stack exchange,提问作者Matthew Casey

