You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow新手求助:两层神经网络训练准确率无提升问题

Troubleshooting Your Two-Layer Neural Network Accuracy Issues

Hey there! Let's work through why your TensorFlow two-layer network isn't improving accuracy with the UCI Heart Disease dataset. I’ll break down common pitfalls and fixes based on the code snippet you shared.

First: Fix the Incomplete Forward Pass

Your forwardprop function cuts off at h_2...—we need to finish defining the full network. Since you have three weight matrices (w_1, w_2, w_3), I assume you’re building an input → hidden layer 1 → hidden layer 2 → output structure. Here’s how to complete it:

def forwardprop(X, w_1, w_2, w_3):
    h_1 = tf.nn.sigmoid(tf.matmul(X, w_1))
    h_2 = tf.nn.sigmoid(tf.matmul(h_1, w_2))  # Complete the second hidden layer
    y_pred = tf.nn.sigmoid(tf.matmul(h_2, w_3))  # Output layer for binary classification
    return y_pred

1. Standardize Your Data (Critical for Neural Networks)

The UCI Heart Disease dataset mixes numerical features (like age, cholesterol) and categorical ones (like sex, chest pain type). Neural networks (especially with sigmoid activation) are sensitive to unnormalized input ranges—values that are too large/small can cause gradient vanishing and slow learning.

Add this preprocessing step before splitting your data:

# Load and prepare data
df = pd.read_csv('heart.csv')  # Update with your dataset path
X = df.drop('target', axis=1).values
y = df['target'].values.reshape(-1, 1)  # Reshape labels to match output layer dimensions

# Standardize features to mean=0, std=1
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Split into train/test sets with scaled data
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2, random_state=RANDOM_SEED)

2. Improve Weight Initialization

Your current init_weights uses tf.random_normal(stddev=0.1), which isn’t ideal for sigmoid activation—this can lead to saturated neurons and vanishing gradients. Switch to Xavier initialization, which is designed for sigmoid/tanh layers:

def init_weights(shape):
    return tf.get_variable("weights", shape, initializer=tf.contrib.layers.xavier_initializer())

3. Swap Sigmoid for a Better Activation Function

Sigmoid is prone to gradient vanishing in even shallow networks. Replace it with ReLU (or LeakyReLU for edge cases) to keep gradients flowing:

def forwardprop(X, w_1, w_2, w_3):
    h_1 = tf.nn.relu(tf.matmul(X, w_1))
    h_2 = tf.nn.relu(tf.matmul(h_1, w_2))
    y_pred = tf.nn.sigmoid(tf.matmul(h_2, w_3))  # Keep sigmoid for binary output
    return y_pred

4. Use the Right Loss Function & Optimizer

For binary classification (heart disease presence/absence), avoid mean squared error. Use binary cross-entropy—and for better numerical stability, pass the raw logits (pre-sigmoid output) to the loss function:

# Define placeholders
input_dim = X_train.shape[1]  # 13 features for UCI Heart Disease
hidden_dim1 = 64  # Adjust based on your needs
hidden_dim2 = 32
output_dim = 1

X = tf.placeholder(tf.float32, shape=[None, input_dim])
y = tf.placeholder(tf.float32, shape=[None, output_dim])

# Initialize weights
w_1 = init_weights((input_dim, hidden_dim1))
w_2 = init_weights((hidden_dim1, hidden_dim2))
w_3 = init_weights((hidden_dim2, output_dim))

# Forward pass (get logits for loss calculation)
h_1 = tf.nn.relu(tf.matmul(X, w_1))
h_2 = tf.nn.relu(tf.matmul(h_1, w_2))
logits = tf.matmul(h_2, w_3)
y_pred = tf.nn.sigmoid(logits)

# Loss function: Binary cross-entropy with logits
loss = tf.reduce_mean(tf.nn.sigmoid_cross_entropy_with_logits(labels=y, logits=logits))

# Use Adam optimizer instead of vanilla SGD (adapts learning rate automatically)
optimizer = tf.train.AdamOptimizer(learning_rate=0.001).minimize(loss)

# Accuracy metric
correct_pred = tf.equal(tf.round(y_pred), y)
accuracy = tf.reduce_mean(tf.cast(correct_pred, tf.float32))

5. Tune Training Parameters

  • Epochs: If you’re only training for a few dozen epochs, your model might not have enough time to learn. Try 500–1000 epochs.
  • Batch Size: Instead of training on the full dataset every time, use mini-batches (e.g., 32 or 64 samples) to stabilize learning:
epochs = 1000
batch_size = 32

with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    
    for epoch in range(epochs):
        # Train in batches
        for i in range(0, len(X_train), batch_size):
            batch_X = X_train[i:i+batch_size]
            batch_y = y_train[i:i+batch_size]
            sess.run(optimizer, feed_dict={X: batch_X, y: batch_y})
        
        # Print progress every 10 epochs
        if epoch % 10 == 0:
            train_acc = sess.run(accuracy, feed_dict={X: X_train, y: y_train})
            test_acc = sess.run(accuracy, feed_dict={X: X_test, y: y_test})
            print(f"Epoch {epoch}: Train Acc = {train_acc:.4f}, Test Acc = {test_acc:.4f}")

6. Check for Class Imbalance

The UCI Heart Disease dataset has a roughly 50/50 split, but if your version is imbalanced, your model might favor the majority class. Fix this with weighted loss:

# Calculate class weights
pos_count = np.sum(y_train == 1)
neg_count = len(y_train) - pos_count
pos_weight = neg_count / pos_count  # Weight for positive class

# Use weighted cross-entropy
loss = tf.reduce_mean(tf.nn.weighted_cross_entropy_with_logits(labels=y, logits=logits, pos_weight=pos_weight))

Final Notes

  • Adjust hidden layer sizes (try 32, 64, or 128) to find a balance between underfitting and overfitting.
  • Add dropout if you see overfitting (train accuracy high, test accuracy low):
h_1 = tf.nn.dropout(tf.nn.relu(tf.matmul(X, w_1)), keep_prob=0.8)
h_2 = tf.nn.dropout(tf.nn.relu(tf.matmul(h_1, w_2)), keep_prob=0.8)

内容的提问来源于stack exchange,提问作者Matthew Casey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:22:30