如何在CNN各层添加0.001的L2权重衰减(TensorFlow实现)
Hey there! I get it—implementing that per-layer L2 weight decay from your sound classification paper can feel a bit tricky at first, but let's walk through it step by step using tf.nn.l2_loss just like you found in that related question.
Here's the core idea:
L2 weight decay works by penalizing large weights to prevent overfitting. The paper specifies a coefficient of 0.001, so we'll calculate a small penalty term for each layer's trainable weights and add all those penalties to our main loss function.
Step-by-Step Implementation
First, let's start by initializing a list to keep track of all our weight decay penalty terms. This makes it easy to sum them up later:
weight_decay_penalties = []
For Convolutional Layers
When you define a convolutional layer (whether using Keras high-level API or lower-level TensorFlow ops), you can grab the layer's kernel weights and compute its decay term:
# Example: Keras Conv2D layer conv_layer = tf.keras.layers.Conv2D( filters=64, kernel_size=(3, 3), activation='relu', padding='same' )(input_features) # Get the trainable kernel weights of the layer conv_weights = conv_layer.kernel # Calculate the L2 decay term for this layer (0.001 * (1/2)*sum(w²)) conv_decay = 0.001 * tf.nn.l2_loss(conv_weights) # Add this penalty to our list weight_decay_penalties.append(conv_decay)
For Dense (Fully Connected) Layers
The process is almost identical for dense layers—we just target the layer's kernel weights (skip the bias terms, since L2 decay is rarely applied to biases):
# Example: Keras Dense layer dense_layer = tf.keras.layers.Dense( units=128, activation='relu' )(flattened_features) # Get the dense layer's trainable weights dense_weights = dense_layer.kernel # Calculate and add the decay term dense_decay = 0.001 * tf.nn.l2_loss(dense_weights) weight_decay_penalties.append(dense_decay)
If You're Using Lower-Level TensorFlow Ops
If you're defining weights directly with tf.Variable (instead of Keras layers), the approach is just as straightforward:
# Define a convolutional weight variable manually conv_weights = tf.Variable( tf.random.truncated_normal([3, 3, input_channels, 64], stddev=0.1), trainable=True ) # Calculate decay term and add to our list conv_decay = 0.001 * tf.nn.l2_loss(conv_weights) weight_decay_penalties.append(conv_decay) # Perform the convolution operation conv_output = tf.nn.conv2d(input_features, conv_weights, strides=[1,1,1,1], padding='SAME')
Combine Penalties with Your Main Loss
Finally, sum all the weight decay penalties and add them to your primary loss (like cross-entropy for classification):
# Assume your main classification loss is stored in `classification_loss` total_loss = classification_loss + tf.math.add_n(weight_decay_penalties)
Quick Note on tf.nn.l2_loss
Just to clarify: tf.nn.l2_loss computes 0.5 * sum(w²) for the given weights. Multiplying this by your 0.001 coefficient gives you exactly the L2 weight decay term specified in the paper—this matches the standard formula for L2 regularization ((λ/2)*sum(w²) where λ=0.001).
内容的提问来源于stack exchange,提问作者Beginner

