You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 2.0不依赖Keras实现多层感知机的技术咨询

Answer

Great question! Let's break this down step by step—first, we'll walk through implementing your MLP without Keras using TensorFlow 2.x's eager execution and GradientTape, then cover how to deepen your understanding, and finally share a few Stack Overflow formatting tips since this is your first post.

Implementing the MLP Without Keras

Your Keras code gives us a clear blueprint—we'll replicate that structure using pure TensorFlow core APIs, focusing on manual layer definition, forward passes, and gradient-based training.

Step 1: Import Dependencies

We'll stick to TensorFlow's core modules (no tf.keras layers or utilities):

import tensorflow as tf
import numpy as np

Step 2: Define Model Parameters

Instead of using Dense layers, we'll manually create trainable variables for weights and biases (mirroring your Keras layer dimensions):

input_dim = 784
hidden_dim = 64
output_dim = 5

# Initialize weights with small random values, biases to 0
W1 = tf.Variable(tf.random.normal([input_dim, hidden_dim], stddev=0.01))
b1 = tf.Variable(tf.zeros([hidden_dim]))

W2 = tf.Variable(tf.random.normal([hidden_dim, hidden_dim], stddev=0.01))
b2 = tf.Variable(tf.zeros([hidden_dim]))

W3 = tf.Variable(tf.random.normal([hidden_dim, output_dim], stddev=0.01))
b3 = tf.Variable(tf.zeros([output_dim]))

Step 3: Forward Pass Function

This function replicates your Keras model's forward flow, including dropout (only active during training):

def forward_pass(x, training=True):
    # First hidden layer: Linear -> ReLU
    x = tf.matmul(x, W1) + b1
    x = tf.nn.relu(x)
    # Apply dropout if training
    if training:
        x = tf.nn.dropout(x, rate=0.5)
    
    # Second hidden layer: Linear -> ReLU
    x = tf.matmul(x, W2) + b2
    x = tf.nn.relu(x)
    if training:
        x = tf.nn.dropout(x, rate=0.5)
    
    # Output layer: Linear -> Softmax
    x = tf.matmul(x, W3) + b3
    return tf.nn.softmax(x)

Step 4: Loss & Evaluation Metrics

We'll implement categorical cross-entropy using a pure TensorFlow approach (no Keras-specific loss wrappers):

def compute_loss(y_pred, y_true):
    # Pure TensorFlow implementation of categorical cross-entropy
    return -tf.reduce_mean(tf.reduce_sum(y_true * tf.math.log(y_pred + 1e-10), axis=1))

def compute_accuracy(y_pred, y_true):
    preds = tf.argmax(y_pred, axis=1)
    truths = tf.argmax(y_true, axis=1)
    return tf.reduce_mean(tf.cast(tf.equal(preds, truths), tf.float32))

Step 5: Prepare Data

Replace tf.keras.utils.to_categorical with TensorFlow's core tf.one_hot function, and use tf.data for efficient batching:

# Assume X (shape: [n_samples, 784]) and y (shape: [n_samples,]) are your raw data
X_train = X[:7000]
y_train = tf.one_hot(y[:7000], depth=output_dim)

X_dev = X[7000:]
y_dev = tf.one_hot(y[7000:], depth=output_dim)

# Use tf.data for efficient batching/shuffling
train_dataset = tf.data.Dataset.from_tensor_slices((X_train, y_train))\
                               .shuffle(len(X_train))\
                               .batch(128)

Step 6: Training Loop with GradientTape

This is where we integrate eager execution and gradient tape to train the model manually:

# Initialize optimizer (TF core optimizer, not Keras-specific)
optimizer = tf.optimizers.SGD(learning_rate=0.01, decay=1e-6, momentum=0.9, nesterov=True)

epochs = 100
for epoch in range(epochs):
    total_loss = 0.0
    total_acc = 0.0
    
    for batch_x, batch_y in train_dataset:
        # Record operations for gradient calculation
        with tf.GradientTape() as tape:
            y_pred = forward_pass(batch_x, training=True)
            loss = compute_loss(y_pred, batch_y)
        
        # Calculate gradients of loss w.r.t. all trainable variables
        grads = tape.gradient(loss, [W1, b1, W2, b2, W3, b3])
        # Update weights using the optimizer
        optimizer.apply_gradients(zip(grads, [W1, b1, W2, b2, W3, b3]))
        
        # Accumulate loss and accuracy for the epoch
        total_loss += loss.numpy() * batch_x.shape[0]
        total_acc += compute_accuracy(y_pred, batch_y).numpy() * batch_x.shape[0]
    
    # Print epoch stats
    avg_loss = total_loss / len(X_train)
    avg_acc = total_acc / len(X_train)
    print(f"Epoch {epoch+1}/{epochs} | Loss: {avg_loss:.4f} | Accuracy: {avg_acc:.4f}")

# Evaluate on dev set
dev_preds = forward_pass(X_dev, training=False)
dev_loss = compute_loss(dev_preds, y_dev)
dev_acc = compute_accuracy(dev_preds, y_dev)
print(f"\nDev Set Performance | Loss: {dev_loss.numpy():.4f} | Accuracy: {dev_acc.numpy():.4f}")

How to Deepen Your Understanding

Here are actionable steps to solidify your grasp of TensorFlow 2.x's core mechanics:

  • Master Eager Execution: Start with TensorFlow's official guides on eager mode to understand how tensors work outside of graphs—focus on immediate evaluation and dynamic computation.
  • Experiment with GradientTape: Try modifying the code to use tape.watch() manually (for non-Variable tensors), test persistent=True for multiple gradient calculations, and debug gradients by printing their values.
  • Build Custom Layers from Scratch: Implement other layer types (like convolutional layers) manually to understand the math behind them, then compare to their Keras equivalents.
  • Dive into Optimizer Logic: Try writing a simple SGD optimizer from scratch (without using tf.optimizers) to see how gradient updates work at the lowest level.
  • Break Down Keras Layers: Inspect the source code of Keras Dense or Dropout layers to see how they encapsulate the manual logic we wrote—this will clarify how Keras abstracts complexity.

Stack Overflow Formatting Tips for Your First Post

Your current draft is already good, but here are a few tweaks to make it more Stack Overflow-friendly:

  • Wrap code in language-specific blocks: Use ```python at the start of code sections to enable syntax highlighting (you already did this, which is great!).
  • Be specific about your confusion: Instead of "不知如何整合", try phrasing it as "I'm unsure how to use GradientTape to compute gradients for a multi-layer network, or how to manually update weights using those gradients."
  • Add context about your data: Briefly mention the shape of your X/y data (e.g., "X is a (n_samples, 784) array of flattened MNIST-like images, y is a (n_samples,) array of class indices") to help others reproduce your setup.
  • Structure your question clearly: Split it into sections like "What I've Done", "What I Need", and "My Confusion" to make it easy to follow.

内容的提问来源于stack exchange,提问作者Anders Stendevad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:45:44