多分类任务中Sigmoid激活的适用性?从零实现神经网络遇Softmax难题
Hey there, let’s break this down step by step—Softmax + cross-entropy can feel tricky at first, but once you get the intuition behind the derivatives, it clicks! I’ll walk you through implementation, common pitfalls, and the magical simplified derivative that makes backprop way easier.
First, the Softmax formula for a single sample’s logits ( z ) is:
[ \sigma(z)_i = \frac{e{z_i}}{\sum_{j=1}C e^{z_j}} ]
where ( C ) is the number of classes.
The big problem here is numerical overflow: if any ( z_i ) is large, ( e^{z_i} ) becomes huge and can turn into infinity. The fix? Subtract the maximum logit from all elements before taking the exponential—this doesn’t change the output, but keeps values manageable:
import numpy as np def softmax(logits): # Subtract max logit per sample to avoid overflow shifted_logits = logits - np.max(logits, axis=1, keepdims=True) exp_values = np.exp(shifted_logits) return exp_values / np.sum(exp_values, axis=1, keepdims=True)
Test this with some extreme values (like logits = [[1000, 2000]])—you’ll get valid probabilities instead of NaNs.
Cross-entropy measures how far your predicted probabilities are from the true one-hot labels. For one-hot encoded ( y_{\text{true}} ) (where only the correct class is 1) and Softmax output ( y_{\text{pred}} ), the loss is:
[ L = -\frac{1}{N} \sum_{i=1}^N \sum_{j=1}^C y_{\text{true},i,j} \cdot \log(y_{\text{pred},i,j}) ]
We add a tiny epsilon to avoid ( \log(0) ) which would give NaNs:
def cross_entropy_loss(y_pred, y_true): epsilon = 1e-10 # Prevent log(0) # Average loss across all samples return -np.sum(y_true * np.log(y_pred + epsilon)) / y_pred.shape[0]
If your labels are class indices (e.g., [0, 2, 1] instead of one-hot), you can adjust the loss function to index directly into the predicted probabilities:
def cross_entropy_loss_from_indices(y_pred, y_true_indices): epsilon = 1e-10 # Get the log probability of the correct class for each sample correct_log_probs = np.log(y_pred[np.arange(len(y_true_indices)), y_true_indices] + epsilon) return -np.mean(correct_log_probs)
Here’s the magic: instead of computing the derivative of Softmax and cross-entropy separately (which involves messy Jacobian matrices), combining them simplifies the gradient drastically.
The gradient of the loss with respect to the input logits (before Softmax) is simply:
[ \frac{dL}{dz} = y_{\text{pred}} - y_{\text{true}} ]
(divided by the number of samples if you’re averaging the loss).
Why? The math simplifies because the cross-entropy loss’s derivative cancels out most of the complexity in Softmax’s derivative. You don’t need to compute the full Jacobian—this formula works for every sample.
Implement this gradient function:
def softmax_cross_entropy_gradient(y_pred, y_true): # y_pred is Softmax output, y_true is one-hot encoded return (y_pred - y_true) / y_pred.shape[0]
If you’re using class indices instead of one-hot labels, first convert them to one-hot:
def indices_to_one_hot(indices, num_classes): return np.eye(num_classes)[indices] # Example: y_true_indices = [0, 2, 1], num_classes=3 # becomes [[1,0,0], [0,0,1], [0,1,0]]
Let’s walk through a quick forward/backward pass example for a final dense layer:
# Initialize weights (input features: 10, classes: 3) W = np.random.randn(10, 3) * 0.01 # Small initial weights to avoid saturation b = np.zeros(3) # Sample input data (5 samples, 10 features) X = np.random.randn(5, 10) # True labels (one-hot encoded) y_true = indices_to_one_hot(np.random.randint(0, 3, 5), num_classes=3) # Forward Pass logits = X @ W + b # Output of dense layer (before Softmax) y_pred = softmax(logits) loss = cross_entropy_loss(y_pred, y_true) # Backward Pass # Gradient of loss w.r.t logits d_logits = softmax_cross_entropy_gradient(y_pred, y_true) # Gradient of loss w.r.t weights W d_W = X.T @ d_logits # Gradient of loss w.r.t biases b d_b = np.sum(d_logits, axis=0) # Update weights (using gradient descent) learning_rate = 0.1 W -= learning_rate * d_W b -= learning_rate * d_b
- Ignoring numerical stability: Always subtract the max logit in Softmax—otherwise you’ll get NaNs/Infs with large logits.
- Confusing logits and Softmax outputs: The gradient formula ( y_{\text{pred}} - y_{\text{true}} ) applies to the logits, not the Softmax output. Don’t waste time computing gradients for the Softmax output itself.
- Forgetting to average loss/gradients: Dividing by the number of samples ensures your loss and gradients scale correctly regardless of batch size.
- Mismatched label formats: Make sure your true labels are either one-hot encoded or you convert indices to one-hot before computing gradients.
Once you get this working, you can plug it into your existing neural network framework—just replace the binary classification activation/loss with this Softmax + cross-entropy setup for multi-class tasks.
内容的提问来源于stack exchange,提问作者KOB

